You will partner closely with engineering, product, and business teams to modernize our statistical tooling, improve self-service experimentation, and extend our measurement framework to emerging AI use cases including LLM evals, prompt evaluation, hybrid human/LLM judging, and offline-to-online quality measurement. Lead the design of scalable statistical frameworks for online experiments across product, business, and operational use cases, including guardrails, heterogeneity analysis, sequential decisioning, variance reduction, and quasi-experimental methods when randomized tests are not feasible.