Hire, grow, and build the culture of a team spanning test engineering, applied ML, and platform-specific expertise Own leadership reporting and safety sign-off representation for Quality Platform Engineering - translating pipeline health, coverage, and findings into risk and product terms for non-technical stakeholders Set and continuously reprioritize which feature clusters, platforms, and hardware form factors the team covers, based on risk, launch timing, and known coverage gaps Drive investment in test pipeline efficiency - turnaround time, infra cost, and flakiness - so pre-merge signal stays fast enough to be useful in the dev loop Extend safety testing coverage beyond iPhone/iOS to every platform and device category Apple ships, partnering with platform teams to close hardware-specific gaps Upstream trusted datasets for product-critical safety cases, so component teams have reliable ground truth to test against early Champion shift-left safety testing - catching issues at the component level before they reach end-to-end evaluation Stay technically fluent enough to evaluate findings, unblock the team on hard problems, and track how model and pipeline behavior shifts across releases Ensure the teams signal stays consistent with Product Evaluations & Researchs ship-time metrics, so pre-merge results reliably predict end-to-end outcomes rather than drifting from them Manage occasional exposure to sensitive or policy-relevant content surfaced during testing, and support team wellbeing around that exposure5+ years of technical team management or leadership experience Experience with ML evaluation, production ML systems, or test/CI infrastructure at scale Strong engineering skills and experience writing production-quality code (Python or similar) Experience working across multiple platforms or hardware form factors, or a demonstrated ability to ramp quickly across unfamiliar platforms Experience working with human-labeled or crowd-sourced evaluation data, including reasoning about label noise and inter-rater agreementExperience working on Responsible AI, AI safety, or trust & safety-adjacent engineering Experience with generative model evaluation and common failure modesStrong organizational and operational skills working with large, multi-functional, diverse teams MS or PhD in Computer Science, Machine Learning, Statistics, or related field, or equivalent experience Familiarity with hardware/platform-specific testing considerations (e.g., on-device constraints, new form factors). Most of this roles impact will come from two things: building a strong team and a healthy team culture from the ground up, and establishing the flywheel that connects Quality Platform Engineering to Product Evaluations & Research and Post-Ship Insights - so fast signal, coverage gaps, and pipeline improvements consistently translate into real engineering decisions and clear leadership visibility, rather than one-off dashboards nobody acts on.