You'll work on voice models, agent capability evals, benchmark design, LLM-as-judge systems, failure analysis, and the infrastructure that ties it together by establishing shared metrics, test sets, and tooling to measure accuracy, resolution quality, and safety consistently across products. Train and fine-tune voice models (TTS, ASR, speech-to-speech) using production and synthetic data, iterating on architecture, data mix, and training strategy to improve accuracy, naturalness, and latency.