Applied AI Researcher

Morpheus Talent Solutions
  • San Mateo, CA
  • Remote
  • $200,000–$350,000 Per Year
Today

Job Description

Job Description

Applied AI Researcher - Model Evaluation & Data Strategy

\n


\n

San Francisco (in-person preferred; open to remote across US, UK, Australia, and Europe) · Retained search - confidential client

\n


\n

The engagement

\n

Morpheus has been exclusively retained to lead the search for a founding Applied AI Researcher on behalf of an early-stage, profitable AI research company.

\n


\n

About the client

\n


\n

An early-stage AI research company that works with leading AI labs to find where frontier models fail and build the expert human data that fixes them. They run a vetted network of 5,000+ top-1% specialists across finance, medicine, law, engineering, music, and other domains - competing on the quality of expert judgment, not the scale of cheap labeling.

\n


\n

Backed by a top pre-seed fund and angel investors who are founders and senior researchers at leading frontier AI labs. Already profitable.

\n


\n

The role

\n


\n

A founding, research-first seat - a genuine thought partner on evaluation and data strategy, not someone who coordinates other people's research, and not client-facing or delivery. You'll design evaluations, form and test hypotheses about model failure, and define the datasets, rubrics, reward signals, and quality controls that move performance - then work with engineers to turn them into scalable programs. Publishing is core to this role, not a perk - you'll be expected to author and present research that positions the company as a research partner to the field, so a prior publication record is essential.

\n


\n

What you'll do

\n


\n
  • \n
  • Design evaluations for generative, reasoning, tool-use, and agentic systems across modalities - text, audio, vision - where technique and domain expertise differ meaningfully by modality.
  • \n
  • Own real experimental design: hypothesize where a model breaks, build the eval to test it, and quantify what's actually failing.
  • \n
  • Build RL environments and reward signals in close partnership with engineers.
  • \n
  • Recommend SFT data, preference data, expert demonstrations, critiques, and eval sets.
  • \n
  • Build quality systems - calibration, blind review, adjudication - that hold up in non-deterministic domains.
  • \n
  • Run pilots that prove whether an intervention moves performance, and publish work that positions the company as a research partner to the field.
  • \n
\n


\n

What the client is looking for

\n


\n
  • \n
  • A track record of published research - you've authored papers at venues like NeurIPS, ICML, ICLR, ACL, or EMNLP (or comparable). This is a research seat with a mandate to publish, so a demonstrated publication record is essential.
  • \n
  • Genuine experimental-design experience - you've designed studies and evals, not just executed someone else's rubric.
  • \n
  • Comfort operating in non-deterministic domains and with novel, fast-moving research frameworks.
  • \n
  • Agentic evaluation proficiency; RL environment experience, ideally built alongside engineers.
  • \n
  • Cross-modal understanding - awareness that audio, text, and vision each demand different techniques.
  • \n
  • Strong Python, model APIs, and structured datasets; solid grounding in benchmark design, human eval, rubric development, and statistical analysis.
  • \n
  • Familiarity with SFT, preference optimization, RLHF/RLAIF, reward modeling, synthetic data, or LLM-as-a-judge.
  • \n
  • Strong technical writing and the ability to drive ambiguous research independently.
  • \n
\n


\n

Nice to have

\n


\n

Experience at an AI lab, foundation-model company, or post-training team; expert-data or human-eval program design; multimodal/coding/agentic eval work; public benchmarks or eval frameworks.

\n


\n

The reality - worth knowing up front

\n


\n

This is an early, high-momentum team that currently works a six-day week: Saturdays are fully remote and self-directed, no set hours - most people use them as a heads-down research day. Compensation is $200K-$350K base + equity; visa sponsorship available.

\n


\n

To apply

\n


\n

Apply here or message me directly. I represent this search exclusively and will share the company name, team, and full details confidentially with candidates who are a strong fit.

Numbers & Facts

LocationSan Mateo, CA (
Remote
)
Salary$200,000–$350,000 Per Year

Skills

  • Adjudicationunmatched
  • Application Programming Interface (API)unmatched
  • Artificial Intelligence (AI)unmatched
  • Benchmarkingunmatched
  • Calibrationunmatched
  • Customer Relationsunmatched
  • Data Modelingunmatched
  • Design Evaluationunmatched
  • Design Servicesunmatched
  • Experiment Designunmatched
  • Financeunmatched
  • Legalunmatched
  • Medicineunmatched
  • Modalityunmatched
  • Musicunmatched
  • Product Demonstrationunmatched
  • Publicationsunmatched
  • Python Programming/Scripting Languageunmatched
  • Research Skillsunmatched
  • Statisticsunmatched
  • Technical Writingunmatched
  • Testingunmatched
  • Training Data Setsunmatched
  • Writing Skillsunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder