ICLR, NeurIPS, ICML, ACL, CVPR, ICCV, FAccT) with a track record in evaluation, alignment, or AI safety Experience with large-scale distributed training (hundreds/thousands of GPUs) and evaluating models in-flight during training Experience with adversarial evaluation and red-teaming, including automated attack generation and jailbreak robustness measurement experience with observability, monitoring, or experiment-tracking systems PhD in Computer Science, Machine Learning, or a relevant technical field Background in statistics and experimental designMeta builds technologies that help people connect, find communities, and grow businesses. This role owns that ground truth: designing the evaluations that detect emerging risks in text, image, voice, video, and agentic systems; building the infrastructure that runs them continuously against training checkpoints and production traffic; and setting the technical direction for how safety is measured across Meta's AI portfolio.