Remote

24-MAG
  • New York, New York
  • Remote
  • $55–$85 Per Hour
30+ days ago

Job Description

We are sharing a specialised full-time consulting opportunity for experienced QA and test engineers with strong expertise in test-case design, end-to-end debugging, quality assurance, Python, and complex technical evaluation workflows.

This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected professionals will review complex multi-step tasks, test reference solutions, identify ambiguity and grading gaps, debug technical environments, and develop repeatable quality processes that keep benchmark results accurate and trustworthy.

Key Responsibilities

Test Case Design

  • Create comprehensive test cases confirming that benchmark tasks function as intended
  • Design positive, negative, boundary, and edge-case tests
  • Validate task requirements, expected outputs, reference solutions, and grading logic
  • Identify scenarios that may produce incorrect or misleading evaluation results
  • Ensure tests measure the intended technical capability accurately

Benchmark Task Review

  • Review complex multi-step tasks and reference solutions before finalisation
  • Identify ambiguous instructions, inconsistent requirements, missing assumptions, and incomplete acceptance criteria
  • Run tasks independently to confirm reproducibility and expected behaviour
  • Assess whether grading standards are clear, fair, and technically defensible
  • Provide actionable feedback to task authors and researchers

Technical Debugging

  • Investigate failures across Python scripts, test harnesses, repositories, and task environments
  • Diagnose unexpected behaviour within unfamiliar codebases
  • Reproduce reported issues and isolate their underlying causes
  • Correct or document environment, dependency, logic, and validation problems
  • Use Git-based workflows to support structured review and collaboration

Quality Process Development

  • Develop practical checklists and repeatable review procedures for benchmark quality
  • Improve consistency across task validation, testing, and approval workflows
  • Document findings clearly so authors can resolve issues efficiently
  • Track recurring defects and recommend preventive quality measures
  • Collaborate closely with researchers, task authors, and other technical reviewers

Benchmark Integrity & Shortcut Detection

  • Examine AI agent runs for unintended shortcuts, loopholes, and grading weaknesses
  • Identify cases where models can receive credit without completing the intended reasoning or technical work
  • Test whether benchmark tasks remain robust across alternative approaches
  • Strengthen evaluation criteria to maintain reliable and meaningful benchmark scores
  • Distinguish valid solution diversity from unintended task exploitation

Ideal Profile

Strong candidates may have:

  • At least 1 year of experience in test engineering, quality assurance, software engineering, research engineering, or a related technical role
  • Demonstrated experience designing test cases and quality-review processes
  • Strong end-to-end debugging skills across complex technical systems
  • Working proficiency in Python and Git
  • Comfort navigating unfamiliar codebases, repositories, and execution environments
  • Exceptional attention to detail and strong written documentation habits
  • Ability to identify ambiguity, edge cases, hidden assumptions, and quality gaps
  • Capacity to work independently through open-ended technical problems
  • Reliable availability for approximately 35 hours per week

Educational Background

  • A master's degree or PhD in a STEM field is highly relevant
  • Equivalent practical experience in an engineering-intensive or research-intensive domain may also be considered
  • Academic or professional experience involving computer science, software engineering, machine learning, mathematics, statistics, or a related technical field may strengthen an application
  • Technical research, open-source contributions, testing projects, or substantial engineering work may also be valuable

Nice to Have

  • Experience with AI training, model evaluation, or quality review of AI-generated work
  • Familiarity with agentic systems and multi-step AI benchmarks
  • Background testing machine learning, research, or data-processing workflows
  • Experience developing automated test suites or validation scripts
  • Familiarity with CI/CD systems, test harnesses, containers, or reproducible environments
  • Experience reviewing reference solutions, grading logic, or technical rubrics
  • Knowledge of adversarial testing, failure-mode analysis, or benchmark design
  • Prior collaboration with AI research or evaluation teams

Why This Opportunity

  • Serve as the quality backbone for advanced agentic AI benchmarks
  • Ensure complex evaluation tasks are accurate, reproducible, and resistant to shortcuts
  • Apply testing and debugging expertise to technically challenging AI research workflows
  • Work closely with researchers and task authors on benchmark improvement
  • Help protect the reliability of evaluation results for frontier AI systems
  • Participate in a structured full-time remote role with competitive hourly compensation

Contract Details

  • Full-time W-2 contingent employment opportunity
  • Fully remote within the United States
  • Expected commitment of approximately 35 hours per week
  • Competitive rates between $55–$85 per hour depending on expertise and project scope
  • Individual task reviews may require one to two days of focused technical work
  • Work may include test design, benchmark review, Python debugging, quality-process development, and shortcut detection
  • Close collaboration with research and task-development teams
  • Engagement scope and duration may evolve according to project requirements and performance

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.

Numbers & Facts

LocationNew York, New York (
Remote
)
Salary$55–$85 Per Hour

Skills

  • Artificial Intelligence (AI)unmatched
  • Benchmarkingunmatched
  • Bug Tracking/Defect Managementunmatched
  • Computer Scienceunmatched
  • Consultingunmatched
  • Continuous Deployment/Deliveryunmatched
  • Continuous Integrationunmatched
  • Debugging Skillsunmatched
  • Detail Orientedunmatched
  • Diversityunmatched
  • Documentationunmatched
  • Failure Analysisunmatched
  • Gitunmatched
  • Identify Issuesunmatched
  • Machine Learningunmatched
  • Mathematicsunmatched
  • Open Sourceunmatched
  • Problem Solving Skillsunmatched
  • Process Developmentunmatched
  • Project Evaluationunmatched
  • Python Programming/Scripting Languageunmatched
  • Quality Assuranceunmatched
  • Quality Assurance Softwareunmatched
  • Quality Metricsunmatched
  • Reliability Analysisunmatched
  • Requirements Validation/Verificationunmatched
  • Scripting (Scripting Languages)unmatched
  • Software Engineeringunmatched
  • Statisticsunmatched
  • System Testunmatched
  • Technical Analysisunmatched
  • Technical Consultingunmatched
  • Technical Researchunmatched
  • Test Automationunmatched
  • Test Caseunmatched
  • Test Designunmatched
  • Test Harnessunmatched
  • Test Plan/Scheduleunmatched
  • Test Suiteunmatched
  • Testingunmatched
  • Validation Testingunmatched
  • Workflow Analysisunmatched
  • Writing Skillsunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder