Remote | QA Test Engineer — $55–$85/hour

24-Mag

  • New York, New York
  • 9 days ago
  • Remote
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Artificial Intelligence (AI)unmatched
    • Benchmarkingunmatched
    • Bug Tracking/Defect Managementunmatched
    • Computer Scienceunmatched
    • Consultingunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • Debugging Skillsunmatched
    • Detail Orientedunmatched
    • Diversityunmatched
    • Documentationunmatched
    • Failure Analysisunmatched
    • Gitunmatched
    • Identify Issuesunmatched
    • Machine Learningunmatched
    • Mathematicsunmatched
    • Open Sourceunmatched
    • Problem Solving Skillsunmatched
    • Process Developmentunmatched
    • Project Evaluationunmatched
    • Python Programming/Scripting Languageunmatched
    • Quality Assuranceunmatched
    • Quality Assurance Softwareunmatched
    • Quality Metricsunmatched
    • Reliability Analysisunmatched
    • Requirements Validation/Verificationunmatched
    • Scripting (Scripting Languages)unmatched
    • Software Engineeringunmatched
    • Statisticsunmatched
    • System Testunmatched
    • Technical Analysisunmatched
    • Technical Consultingunmatched
    • Technical Researchunmatched
    • Test Automationunmatched
    • Test Caseunmatched
    • Test Designunmatched
    • Test Harnessunmatched
    • Test Plan/Scheduleunmatched
    • Test Suiteunmatched
    • Testingunmatched
    • Validation Testingunmatched
    • Workflow Analysisunmatched
    • Writing Skillsunmatched

    Description

    We are sharing a specialised full-time consulting opportunity for experienced QA and test engineers with strong expertise in test-case design, end-to-end debugging, quality assurance, Python, and complex technical evaluation workflows.

    This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected professionals will review complex multi-step tasks, test reference solutions, identify ambiguity and grading gaps, debug technical environments, and develop repeatable quality processes that keep benchmark results accurate and trustworthy.

    Key Responsibilities

    Test Case Design

    • Create comprehensive test cases confirming that benchmark tasks function as intended
    • Design positive, negative, boundary, and edge-case tests
    • Validate task requirements, expected outputs, reference solutions, and grading logic
    • Identify scenarios that may produce incorrect or misleading evaluation results
    • Ensure tests measure the intended technical capability accurately

    Benchmark Task Review

    • Review complex multi-step tasks and reference solutions before finalisation
    • Identify ambiguous instructions, inconsistent requirements, missing assumptions, and incomplete acceptance criteria
    • Run tasks independently to confirm reproducibility and expected behaviour
    • Assess whether grading standards are clear, fair, and technically defensible
    • Provide actionable feedback to task authors and researchers

    Technical Debugging

    • Investigate failures across Python scripts, test harnesses, repositories, and task environments
    • Diagnose unexpected behaviour within unfamiliar codebases
    • Reproduce reported issues and isolate their underlying causes
    • Correct or document environment, dependency, logic, and validation problems
    • Use Git-based workflows to support structured review and collaboration

    Quality Process Development

    • Develop practical checklists and repeatable review procedures for benchmark quality
    • Improve consistency across task validation, testing, and approval workflows
    • Document findings clearly so authors can resolve issues efficiently
    • Track recurring defects and recommend preventive quality measures
    • Collaborate closely with researchers, task authors, and other technical reviewers

    Benchmark Integrity & Shortcut Detection

    • Examine AI agent runs for unintended shortcuts, loopholes, and grading weaknesses
    • Identify cases where models can receive credit without completing the intended reasoning or technical work
    • Test whether benchmark tasks remain robust across alternative approaches
    • Strengthen evaluation criteria to maintain reliable and meaningful benchmark scores
    • Distinguish valid solution diversity from unintended task exploitation

    Ideal Profile

    Strong candidates may have:

    • At least 1 year of experience in test engineering, quality assurance, software engineering, research engineering, or a related technical role
    • Demonstrated experience designing test cases and quality-review processes
    • Strong end-to-end debugging skills across complex technical systems
    • Working proficiency in Python and Git
    • Comfort navigating unfamiliar codebases, repositories, and execution environments
    • Exceptional attention to detail and strong written documentation habits
    • Ability to identify ambiguity, edge cases, hidden assumptions, and quality gaps
    • Capacity to work independently through open-ended technical problems
    • Reliable availability for approximately 35 hours per week

    Educational Background

    • A master's degree or PhD in a STEM field is highly relevant
    • Equivalent practical experience in an engineering-intensive or research-intensive domain may also be considered
    • Academic or professional experience involving computer science, software engineering, machine learning, mathematics, statistics, or a related technical field may strengthen an application
    • Technical research, open-source contributions, testing projects, or substantial engineering work may also be valuable

    Nice to Have

    • Experience with AI training, model evaluation, or quality review of AI-generated work
    • Familiarity with agentic systems and multi-step AI benchmarks
    • Background testing machine learning, research, or data-processing workflows
    • Experience developing automated test suites or validation scripts
    • Familiarity with CI/CD systems, test harnesses, containers, or reproducible environments
    • Experience reviewing reference solutions, grading logic, or technical rubrics
    • Knowledge of adversarial testing, failure-mode analysis, or benchmark design
    • Prior collaboration with AI research or evaluation teams

    Why This Opportunity

    • Serve as the quality backbone for advanced agentic AI benchmarks
    • Ensure complex evaluation tasks are accurate, reproducible, and resistant to shortcuts
    • Apply testing and debugging expertise to technically challenging AI research workflows
    • Work closely with researchers and task authors on benchmark improvement
    • Help protect the reliability of evaluation results for frontier AI systems
    • Participate in a structured full-time remote role with competitive hourly compensation

    Contract Details

    • Full-time W-2 contingent employment opportunity
    • Fully remote within the United States
    • Expected commitment of approximately 35 hours per week
    • Competitive rates between $55–$85 per hour depending on expertise and project scope
    • Individual task reviews may require one to two days of focused technical work
    • Work may include test design, benchmark review, Python debugging, quality-process development, and shortcut detection
    • Close collaboration with research and task-development teams
    • Engagement scope and duration may evolve according to project requirements and performance

    About the Platform

    This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

    By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.

    Numbers & Facts

    LocationNew York, New York (
    Remote
    )
    Website4-mag.com/privacy-policy

    Similar Jobs

    See more jobs