We are sharing a specialised part-time consulting opportunity for experienced machine learning engineers with hands-on experience using AI coding agents and building production ML systems, model deployment infrastructure, LLM applications, or AI-powered products.
This sprint-based role supports an advanced AI research initiative focused on evaluating frontier coding models through realistic machine learning engineering workflows. Selected professionals will use AI coding agents to complete technical tasks, review model-generated implementations, identify bugs and failure modes, and compare how different models perform across practical ML engineering scenarios.
Key Responsibilities
Machine Learning Engineering Evaluation
- Review complex machine learning and AI engineering tasks completed with frontier coding agents
- Evaluate implementations involving model training, inference systems, MLOps, and LLM applications
- Assess technical correctness, architecture choices, implementation quality, and engineering trade-offs
- Apply professional ML engineering judgment to realistic production-oriented scenarios
AI Coding Agent Testing
- Use AI coding agents as part of hands-on technical workflows
- Evaluate how effectively coding models interpret requirements and implement solutions
- Identify bugs, incomplete implementations, edge cases, and unexpected behaviour
- Assess where models require additional prompting, correction, or manual engineering intervention
Technical Quality & Failure Analysis
- Identify performance issues, reliability problems, and model failure modes
- Review generated code for maintainability, correctness, and practical usability
- Evaluate whether implementations would function appropriately in realistic ML environments
- Document technical strengths, weaknesses, and important implementation risks
Model Comparison & Technical Judgment
- Compare outputs produced by multiple frontier coding models
- Assess differences in implementation strategy, code quality, technical reasoning, and reliability
- Determine which approaches best satisfy task requirements
- Provide clear written assessments explaining relevant engineering trade-offs
Ideal Profile
Strong candidates may have:
- At least 2 years of professional machine learning engineering experience
- Experience building production ML systems, AI-powered applications, or model-serving infrastructure
- Hands-on experience with model training, inference, deployment, or MLOps
- Experience developing LLM applications or integrating foundation models into production systems
- Regular use of AI coding agents within software or machine learning development workflows
- Strong ability to evaluate model-generated code and technical implementation decisions
- Excellent debugging, analytical reasoning, and written communication skills
- Ability to work efficiently within short, intensive project sprints
Educational Background
- A degree in computer science, machine learning, artificial intelligence, software engineering, or a related technical discipline may be helpful
- Advanced study in machine learning or computer science may strengthen an application
- Equivalent professional experience building and deploying production ML systems may also be considered
- Practical engineering depth is particularly important for this engagement
Nice to Have
- Experience with Cursor, Claude Code, Codex, Windsurf, Gemini CLI, or comparable AI coding tools
- Production experience deploying machine learning models
- Familiarity with model-serving architectures and inference optimisation
- Experience building LLM-powered applications or agentic systems
- Knowledge of MLOps, deployment pipelines, monitoring, or model infrastructure
- Experience evaluating generated code across multiple AI coding systems
- Previous exposure to AI evaluation, benchmark development, or structured technical review
Why This Opportunity
- Work directly with frontier AI coding agents on realistic ML engineering problems
- Evaluate advanced models across production-oriented machine learning workflows
- Apply practical engineering experience to identify subtle technical failure modes
- Compare multiple coding systems and help improve their reliability
- Participate in intensive technical sprints with task-based compensation
Contract Details
- Independent contractor role
- Fully remote with flexible scheduling
- Sprint-based project with task windows typically spanning approximately 12–24 hours
- Compensation is $400 per accepted task
- Typical tasks require approximately 2–3 hours after ramp-up
- Compensation is tied to successfully accepted work
- Work may include ML implementation review, coding-agent evaluation, debugging, model comparison, and technical analysis
- Weekly payments via Stripe or Wise
- Projects may be extended, shortened, or adjusted depending on scope and performance
- Work will not involve access to confidential or proprietary information from any employer, client, or institution
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.