Research Engineer, Benchmarks

Clera
  • San Francisco, California
  • $150,000–$250,000 Per Year
2 days ago

Job Description

About the Role

This is a high-ownership research engineering role on a small, technical team focused on designing and building benchmarks that evaluate frontier AI agents on realistic, domain-specific workflows. You'll work alongside researchers and engineers to ensure evaluations are rigorous, credible, and trusted by AI labs and enterprise customers. The quality of these benchmarks directly shapes how the world measures and improves AI agent performance.

What You'll Do

  • Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.

  • Partner with subject-matter experts to define realistic workflows and translate them into evaluation tasks and criteria.

  • Build and operate reliable infrastructure to run models and agents against benchmark tasks at scale.

  • Develop metrics and statistical analyses to measure benchmark difficulty, reliability, and failure modes.

  • Validate that benchmark performance correlates with real-world evaluations and frontier lab expectations.

  • Write clear technical documentation and benchmark reports for research and engineering audiences.

What We're Looking For

  • 2–4 years of experience in research engineering, ML engineering, or related roles — with a focus on AI benchmarks, evaluation infrastructure, or agent environments.

  • Strong proficiency in Python, Docker, and Linux environments for building research or production infrastructure.

  • Experience designing and running benchmarks or evaluation environments for AI agents or large language models.

  • Experience developing metrics, statistical analyses, or validation studies to assess benchmark quality and real-world correlation.

  • Ability to collaborate with domain experts to translate complex workflows into well-scoped evaluation tasks.

  • Deep intuition for what makes a benchmark realistic, reliable, and practically useful.

  • Strong attention to detail with a habit of catching subtle inconsistencies and edge cases.

  • Comfort working independently in fast-moving, unstructured, early-stage environments.

  • Excellent written communication skills; able to make technical results legible across time zones and teams.

  • Nice to have: Published papers or technical writing on AI benchmarking, model evaluation, or failure modes; experience with RL training pipelines or data generation; background at frontier AI labs, research institutions, or on widely used public benchmark projects.

Compensation & Benefits

Salary range: $150,000 – $250,000 USD annually. Equity offered. Visa sponsorship is available.

Location

On-site in San Francisco, CA, USA. In-person presence is expected.

Numbers & Facts

LocationSan Francisco, California
Salary$150,000–$250,000 Per Year

Skills

  • Artificial Intelligence (AI)unmatched
  • Artificial Intelligence (AI) Agentsunmatched
  • Benchmarkingunmatched
  • Communication Skillsunmatched
  • Detail Orientedunmatched
  • Dockerunmatched
  • Environmental Researchunmatched
  • Linux Operating Systemunmatched
  • Metricsunmatched
  • Modeling Languagesunmatched
  • Performance Managementunmatched
  • Publicationsunmatched
  • Python Programming/Scripting Languageunmatched
  • Research Laboratoryunmatched
  • Statisticsunmatched
  • Team Playerunmatched
  • Technical Publicationsunmatched
  • Technical Writingunmatched
  • Writing Skillsunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder