Evals owners own execution of verification pipelines within their products; this role ensures consistency and identifies gaps across the portfolio while building institutional competence by surfacing performance patterns and proven methodologies, enabling evals captains' ability to execute and unblocking them as needed Defines what leadership needs to see, how model health should be measured and reported, and what thresholds trigger escalation Provides thought partnership to evals managers on narrative of model health, provides visibility into our classification strategy and accuracy measurement process Works with evaluation managers to drive cross-app taxonomy alignment in alignment with cross-functional needs and advises on a strategy for the migration of LLM accuracy assessment to judges Owns the consolidated view of all production model performance, identifies systemic patterns and emerging risks, and ensures leadership can verify model health on demand Partners with AI Implementations, operational systems teams and the Metrics & Measurement team to build and maintain the infrastructure that surfaces this information Establishes performance guardrails that evals captains implement. Continuously scans industry developments and best practices to incorporate into org-wide approach Maintains a tight feedback loop with product and eng teams across apps to ensure alignment on production priorities and deployment risks Deploys deep SME expertise to diagnose, unblock and directly resolve technical bottlenecks to complex model quality problems (atrophy, accuracy regressions, performance plateaus) when evaluation leads encounter blockers they cannot resolve independently Drives alignment with cross-functional teams (quality and reliability partner teams) on tooling needs to support Product Operations classification strategy (ML classification tooling for initial-tier classification, user voice, breakdown graphs).