Remote | Member of Technical Staff, Coding Research — $600,000–$1,300,000/year

24-Mag
  • New York, New York
  • Remote
    6 days ago

    Job Description

    We are sharing a specialised full-time opportunity for experienced technical professionals with strong backgrounds in software engineering, AI research, model evaluation, or machine learning to contribute to the evaluation and development of frontier coding agents.

    Selected professionals will work at the intersection of AI research, software engineering, and model evaluation, designing benchmarks, methodologies, datasets, and technical systems that determine how advanced coding models are measured and improved. The role combines coding-agent evaluation, failure analysis, large-scale experimentation, research tooling, and close collaboration with researchers and engineers.

    Key Responsibilities

    Coding Agent Evaluation & Benchmark Design

    • Design and own evaluation frameworks for advanced coding agents
    • Develop benchmark specifications, scoring methodologies, rubrics, and quality standards
    • Establish rigorous methods for measuring coding-model performance across diverse software-engineering tasks
    • Define objective criteria for correctness, reasoning quality, robustness, and task completion
    • Maintain methodological rigour, reproducibility, and consistency across evaluation workflows

    Datasets, Golden Examples & Evaluation Protocols

    • Develop high-quality datasets, golden examples, and structured evaluation protocols
    • Design technical tasks that enable reliable assessment of frontier coding systems
    • Build data and evaluation workflows that support model development and iterative improvement
    • Identify gaps in benchmark or dataset coverage and develop new evaluation categories where needed
    • Apply strong technical judgement when determining whether evaluation data provides meaningful research signal

    Model Behaviour & Failure Analysis

    • Analyse coding-agent behaviour and identify systematic weaknesses, failure modes, and performance limitations
    • Investigate incorrect reasoning, implementation errors, tool-use failures, and incomplete task execution
    • Translate findings into actionable recommendations for model training and evaluation
    • Design experiments to test hypotheses regarding coding-model capabilities
    • Use evaluation results to guide improvements in datasets, methodologies, and system performance

    Research Tooling & Technical Collaboration

    • Build tooling and infrastructure supporting large-scale experimentation, data generation, review workflows, and evaluation pipelines
    • Automate technical processes to improve evaluation efficiency and research velocity
    • Collaborate closely with researchers, engineers, and applied AI teams
    • Contribute to technical reports, benchmark studies, research documentation, and external-facing research initiatives
    • Communicate complex technical findings clearly to both specialist and broader technical audiences

    Ideal Profile

    • Strong software-engineering background with expertise in Python, C++, or comparable programming languages
    • Minimum of 3 years of experience in software engineering, machine learning, AI research, evaluation, or a related technical discipline
    • Experience designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies
    • Familiarity with large language models, coding agents, reinforcement learning, model evaluation, or related AI systems
    • Proven ability to build tooling, automate workflows, and improve technical processes through systematic experimentation
    • Strong analytical skills and ability to investigate complex model behaviour and technical failure modes
    • Excellent written and verbal communication skills
    • Ability to operate effectively in fast-moving research environments with significant ambiguity and evolving priorities
    • Experience with frontier AI systems, coding agents, or model-evaluation research is advantageous
    • Experience designing benchmarks or datasets for machine-learning systems at scale is strongly valued
    • Familiarity with agentic workflows, tool use, reinforcement learning, or post-training methodologies is beneficial
    • Publications, open-source contributions, or demonstrated technical leadership in AI, machine learning, or software engineering are advantageous
    • Strong interest in understanding how data, evaluations, and feedback mechanisms influence model capabilities

    Engagement Details

    • Full-time engagement
    • Fully remote
    • Compensation: $600,000–$1,300,000/year
    • Work will involve coding-agent evaluation, benchmark development, dataset design, model failure analysis, experimentation, and research tooling
    • Responsibilities will span both research methodology and hands-on technical implementation
    • Collaboration will involve researchers, engineers, applied AI teams, and other technical stakeholders
    • Research priorities, benchmarks, datasets, and evaluation methodologies may evolve as model capabilities develop
    • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party

    About the Platform

    This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

    By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy

    Numbers & Facts

    LocationNew York, New York (
    Remote
    )
    Website4-mag.com/privacy-policy

    Skills

    • Analysis Skillsunmatched
    • Artificial Intelligence (AI)unmatched
    • Benchmarkingunmatched
    • C++ Programming Languageunmatched
    • Communication Skillsunmatched
    • Data Analysisunmatched
    • Data Setsunmatched
    • Documentationunmatched
    • Experiment Designunmatched
    • Failure Analysisunmatched
    • Machine Learningunmatched
    • Machine Toolunmatched
    • Modeling Languagesunmatched
    • Open Sourceunmatched
    • Performance Modelingunmatched
    • Presentation/Verbal Skillsunmatched
    • Process Improvementunmatched
    • Programming Languagesunmatched
    • Project Evaluationunmatched
    • Publicationsunmatched
    • Python Programming/Scripting Languageunmatched
    • Quality Metricsunmatched
    • Reinforcement Learningunmatched
    • Software Engineeringunmatched
    • Systems Analysisunmatched
    • Technical Analysisunmatched
    • Technical Consultingunmatched
    • Technical Leadershipunmatched
    • Technical Researchunmatched
    • Technical/Engineering Designunmatched
    • Test Designunmatched
    • Training Data Setsunmatched
    • Workflow Analysisunmatched
    • Writing Skillsunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder