Sanas is building a full Speech AI suite, all working together as one platform. As that surface area grows, so does the need for a single, rigorous owner of how we know everything is working.
As Evaluations Lead, you'll design the evaluation frameworks and benchmarking systems that answer that question — sitting at the intersection of research, product, and infrastructure to build the metrics, systems, and studies that hold our models accountable. This role suits someone who pairs scientific rigor with real technical execution. Your work will shape how Sanas builds and evaluates its models across every one of these products, making sure progress is measured not just by static benchmarks, but by the harder, more meaningful qualities — understanding, naturalness, and adaptability in real-world interaction.
Your Impact
Identify and define the model capabilities and behaviors that actually matter for evaluation — not just what's easy to measure
Build and ship evaluation pipelines with robust statistical analysis and clear, actionable reporting
Partner directly with model training and research teams to embed evaluation into the development loop itself
Prototype new user studies and behavioral experiments that ground evaluation in how these models actually get used
What You Bring
Experience designing or implementing evaluation frameworks for generative models — audio, text, or multimodal
Strong technical and analytical skills, with the ability to take an open-ended research idea and turn it into a production-ready system
Creativity in defining novel, quantitative metrics for qualities that are inherently subjective
Genuine excitement for building evaluation systems that bridge research and real-world use
Equal parts curious and rigorous — driven by actually figuring out how to measure meaningful progress, not just reporting a number
The ability to build it yourself. This is an engineering role — you'll be writing the pipelines and tooling, not just specifying them
Nice-to-Haves
AI modeling experience — someone who has trained, fine-tuned, or shipped models themselves brings a level of judgment to evaluation design that's hard to substitute, and is highly valued for this role
Background in audio modeling — Speech-to-Text, Text-to-Speech, or similar
Multilingual — especially relevant for evaluating multilingual systems, where understanding the semantic nuance across languages, not just the literal accuracy, is core to getting evaluation right
Numbers & Facts
Location
Palo Alto, California
Skills
Analysis Skillsunmatched
Artificial Intelligence (AI)unmatched
Benchmarkingunmatched
Customer/Consumer Behaviorunmatched
Machine Toolunmatched
Metricsunmatched
Multilingualunmatched
Production Systemsunmatched
Prototypingunmatched
Quality Metricsunmatched
Statisticsunmatched
Technical Researchunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.