The person in this role will help teams build a deep understanding of how agents behave, where they succeed, where they fail, which gaps matter most, and where those gaps should be addressed: in prompts, tools, orchestration, retrieval, ranking, product UX, safety systems, or core code. Demonstrated experience evaluating LLMs, including designing eval datasets, defining quality metrics, analyzing model behavior, identifying failure modes, and using results to guide product or system improvements.