The System Triage Services team builds mission-critical applications that help detect, analyze, and classify low-level software crashes across all Apple platforms. You will design and build the next generation of AI-powered triage solutions, which serves as the first line of defense for all technologies developed by Core OS.
This role blends distributed systems engineering with hands-on agent development, tackling problems in reliability, scale, and system design alongside applied AI. If you want to understand operating systems at a deep level and are motivated by shipping work with measurable, visible impact, this role is for you! You will set the technical direction for agentic systems that change how Core OS triages software bugs at scale. Youll own the architecture of production infrastructure that runs AI/ML workloads reliably and gets better over time through rigorous evaluation and feedback loops. The goal is to make autonomous triage a trusted, first-class part of how Core OS ships and debugs software across every Apple platform.
This is a senior technical leadership role. Youll turn ambiguous, high-stakes triage problems into clear system designs. Youll make the build-vs-reuse and reliability tradeoffs, and youll influence engineering teams across Software, Hardware, and Silicon groups to adopt what you build. Your decisions will shape the platform other teams build on.Architect and lead delivery of the end-to-end agentic triage platform, including agent orchestration, tool integration, data pipelines, and the build and integration lifecycle, with a focus on reliability, latency, and cost at Apple scale. Define the evaluation strategy for production agents: offline benchmarks, golden datasets, regression gates, drift detection, and confidence calibration. Set the quality bar for when automation can act without a human. Build observability so failures, drift, and low-confidence routing decisions surface before they reach engineers, and turn those signals into measurable improvements. Work with Core OS, Hardware, and Silicon engineering leads to find high-leverage automation opportunities, align on requirements, and drive adoption beyond the immediate team. Own production agents through their whole life, from prototype to hardened service, including operational excellence and on-call practices for triage-critical infrastructure. Mentor and grow engineers in agent design, evaluation methodology, and distributed systems. Raise the technical bar through design reviews and hands-on example-setting.6+ years of software engineering experience building and operating production distributed systems or data platforms at scale. Expert-level Python and strong system design skills, with a track record of owning architecture for services other teams depend on. 2+ years building and shipping LLM-based or agentic systems to production, including tool calling, orchestration, and guardrails. Hands-on experience designing evaluation frameworks for ML/LLM systems, such as offline benchmarks, regression testing, drift detection, and confidence scoring. Deep working knowledge of RAG architectures, embedding generation, vector databases, and data preparation for agentic workflows. Proven technical leadership: driving multi-quarter projects through ambiguity, influencing across org boundaries, and mentoring other engineers. BS in Computer Science or a related field, or equivalent practical experience.MS or PhD in Computer Science, Machine Learning, or a related field. Background in kernel or OS-level debugging, crash and panic analysis, or root-cause investigation on complex systems. Experience with large-scale telemetry or observability systems feeding automated decision-making. Track record of shipping platforms or developer tools widely adopted by engineering organizations beyond your own. Experience setting technical strategy and roadmaps, and presenting tradeoffs and outcomes to senior engineering leadership. Excellent written and verbal communication, including writing design documents that align multiple teams.
| Location | Cupertino, CA |
| Industry | Computer/IT Services |
| Company Size | 10,000 employees or more |
| Year Founded | 1976 |
| Website | https://www.apple.com/jobs |
We’re a diverse collection of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. The people who work here have reinvented entire industries with the Mac, iPhone, iPad, and Apple Watch, as well as with services, including iTunes, the App Store, Apple Music, and Apple Pay. And the same passion for innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it.
There’s a place here for every kind of brilliant. Everyone here is an innovator, or an innovator-to-be, no matter what your team or your role. So bring your passion, courage, and original thinking and get ready to share it, because every new product, service, or feature we invent is the result of people working together to make each others’ ideas stronger. Innovation at this level depends on people who represent the variety of the human experience and inspire us with their own fresh perspectives. Together, we’ll do amazing work that can make a difference in people’s lives. Including your own. Learn more about working at Apple.
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.
Free resume templatesImprove your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.
Free resume builder