Apples AIML Evaluation team builds the systems and methodologies that measure and improve the quality of foundation models and agentic experiences. We are looking for a senior, hands-on Machine Learning Engineering Manager to lead a small team working at the intersection of model evaluation, agent optimization, and data generation. In this role, you will help define how evaluation closes the loop with model and product development, turning observed quality gaps into targeted improvements to prompts, agent harnesses, datasets, and models.
You will combine technical depth with people leadership. You should be comfortable moving from research papers and experimental results to production-quality ML pipelines, while mentoring engineers and aligning teams around a clear technical direction. Your work will span Apple Foundation Models and product teams, with the goal of creating repeatable evaluation and refinement loops that improve the quality of Apple intelligence experiences. As a Senior Machine Learning Engineering Manager in AIML Evaluation, you will lead the technical strategy and execution for agent evaluation and automatic optimization. You will own systems that evaluate foundation models and agents, diagnose failure modes, and use those signals to drive automated prompt, context, tool, rubric, and agent-harness improvements. You will also help establish the interfaces between evaluation and post-training so that high-value failures can be converted into targeted data, environments, reward signals, and measurable model improvements.
This is a hands-on leadership role. You will prototype new approaches, participate in architecture and code reviews, design experiments, and help your team translate emerging research into scalable evaluation and optimization pipelines. You will partner closely with Apple Foundation Models, product engineering teams, and other AIML groups to build an evaluation flywheel that connects real product behavior with model and agent refinement. You will also work across the organization to advance synthetic data generation for both evaluation and post-training, with strong attention to data quality, representativeness, privacy, and reproducibility.Architects and builds scalable evaluation systems for foundation models and agents, including benchmarks, LLM-based evaluators, simulation environments, trajectory analysis, and regression testing. Establishes an end-to-end evaluation flywheel with Apple Foundation Models and product teams that connects observed failures to diagnosis, targeted refinement, post-training, and measurable quality improvement. Leads, mentors, and grows a small team of machine learning engineers while remaining deeply involved in technical design, experimentation, implementation, and review. Defines the technical strategy and roadmap for automatic prompt, context, tool, rubric, and agent-harness optimization for agentic development and model evaluation. Develops methods that convert evaluation findings into actionable model-improvement signals, including targeted datasets, synthetic trajectories, reward or preference signals, and optimization objectives. Partners across AIML to design and scale synthetic data generation pipelines for evaluation and post-training. Applies and adapts recent research in LLM and agent evaluation, automatic optimization, LLM-as-judge, reward modeling, test-time search, and post-training to production-quality workflows.8+ years of professional experience in machine learning, applied research, or software engineering, including experience building production ML systems or large-scale experimentation platforms. 3+ years of technical leadership experience, including direct people management of machine learning or software engineers and a demonstrated ability to mentor and grow strong technical talent. Master's or PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related technical field. Strong hands-on programming and software engineering skills, particularly in Python, with experience building reliable ML pipelines using modern machine learning or deep learning frameworks. Deep experience with large language models or agentic systems, including evaluation of multi-turn behavior, tool use, planning, reasoning, or other action-taking workflows. Experience building automated evaluation methods such as LLM-based judges, rubrics, reward models, simulation-based evaluation, or scalable benchmark infrastructure. Experience with at least one model or agent refinement area such as automatic prompt or context optimization, post-training, preference optimization, reinforcement learning, or agent-harness optimization. Excellent communication and collaboration skills, with demonstrated ability to align research, engineering, and product teams around ambiguous technical problems.Track record of applying recent machine learning research to production systems or high-impact product development. Experience with automatic prompt or context optimization, agent-search methods, evaluator optimization, or multi-objective optimization for agentic systems. Experience generating and evaluating synthetic datasets, tool-use trajectories, or multi-turn agent interactions, including methods for filtering, deduplication, diversity, and quality control. Experience designing evaluation systems that combine offline benchmarks, simulation, human evaluation, and product- or usage-derived signals. Experience with privacy-preserving or on-device machine learning and evaluation. Demonstrated ability to influence technical strategy across organizational boundaries and communicate complex model-quality tradeoffs to senior technical leaders.
| Location | Cupertino, CA |
| Industry | Computer/IT Services |
| Company Size | 10,000 employees or more |
| Year Founded | 1976 |
| Website | https://www.apple.com/jobs |
We’re a diverse collection of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. The people who work here have reinvented entire industries with the Mac, iPhone, iPad, and Apple Watch, as well as with services, including iTunes, the App Store, Apple Music, and Apple Pay. And the same passion for innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it.
There’s a place here for every kind of brilliant. Everyone here is an innovator, or an innovator-to-be, no matter what your team or your role. So bring your passion, courage, and original thinking and get ready to share it, because every new product, service, or feature we invent is the result of people working together to make each others’ ideas stronger. Innovation at this level depends on people who represent the variety of the human experience and inspire us with their own fresh perspectives. Together, we’ll do amazing work that can make a difference in people’s lives. Including your own. Learn more about working at Apple.
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.
Free resume templatesImprove your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.
Free resume builder