We are sharing a specialised part-time consulting opportunity for experienced software engineers with strong expertise in Python, Java, Rust, C++, Go, TypeScript, algorithms, debugging, refactoring, and performance optimisation to contribute to an advanced AI training project involving Model Context Protocol (MCP) environments.
Selected professionals will create reinforcement-learning environments that test an AI model's ability to solve complex software-engineering problems using MCP tools and real server interactions. The work combines practical software engineering, deterministic evaluation design, and the creation of high-quality reference solutions. No prior experience in AI is required.
Key Responsibilities
MCP Environment Development
-
Create reinforcement-learning environments based on realistic software-engineering tasks
-
Design scenarios requiring agents to discover and reason over information from MCP servers
-
Build tasks involving real tool interactions rather than isolated code-generation exercises
-
Ensure environments accurately measure both MCP tool use and engineering capability
-
Maintain reproducibility across evaluation runs
Software Engineering Task Design
-
Create challenging tasks involving bug fixing, feature implementation, refactoring, and performance optimisation
-
Develop scenarios that require meaningful reasoning across existing codebases
-
Design tasks that test algorithms, data structures, debugging, and architectural judgement
-
Ensure problems reflect realistic engineering constraints and workflows
-
Balance task complexity with clear, measurable success criteria
Golden Solutions & Deterministic Verification
-
Create high-quality golden reference solutions for evaluation tasks
-
Develop deterministic verification logic that reliably distinguishes correct from incorrect implementations
-
Define clear acceptance criteria for software behaviour and task completion
-
Validate environments against edge cases and unintended solution paths
-
Ensure evaluation logic remains stable and reproducible
Code Quality & Performance Engineering
-
Debug complex software issues across multiple programming languages
-
Implement maintainable features in existing codebases
-
Refactor code while preserving intended functionality
-
Identify and resolve performance bottlenecks
-
Apply scalability, maintainability, and software-quality best practices
Technical Review & Collaboration
-
Review task quality, code correctness, and evaluation robustness
-
Communicate technical decisions and assumptions clearly
-
Participate in collaborative review of software-engineering environments
-
Contribute to code-review standards and engineering best practices
-
Work effectively in remote and cross-functional technical teams
Ideal Profile
-
Strong proficiency in one or more of C++, Python, Java, Go, TypeScript, or Rust
-
Deep understanding of algorithms, data structures, and performance optimisation
-
Demonstrated experience debugging complex software issues
-
Strong background in feature development and codebase refactoring
-
Proven ability to improve software performance and scalability
-
Experience working with large or distributed codebases is highly valuable
-
Familiarity with rigorous code-review practices and software-engineering standards
-
Strong written and verbal communication skills
-
High attention to technical detail and reproducibility
-
Experience with modern AI or machine-learning systems is beneficial but not required
-
Prior AI-training or model-evaluation experience is not required
Engagement Details
-
Part-time independent contractor engagement
-
Fully remote
- Compensation: $60–$120/hour
- Expected commitment: approximately 15 hours per week
-
Schedule is flexible, including the option to work evenings or weekends
-
Compensation is output-based, with payment made for tasks that meet project specifications
-
Minimum weekly submission requirements apply
-
Work will involve MCP-based reinforcement-learning environments, software-engineering task design, deterministic verification, and golden reference solutions
-
The selection process may include screening questions, an approximately 30-minute AI interview, a technical assessment, and hiring-manager review
-
Selected professionals should be prepared to begin their first tasks within approximately 24–48 hours of completing onboarding
-
Roles are typically filled within approximately 48 hours
-
Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy