Apple Inc logo

Machine Learning Engineer - AI & ML Evaluation Frameworks

Apple Inc

  • Cupertino, CA
  • 30+ days ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Algorithmsunmatched
    • Appleunmatched
    • Artificial Intelligence (AI)unmatched
    • Communication Skillsunmatched
    • Computer Scienceunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • Customer/Client Researchunmatched
    • Data Managementunmatched
    • Data Processingunmatched
    • Data Visualizationunmatched
    • Deep Learningunmatched
    • Demographicsunmatched
    • Engineeringunmatched
    • Failure Analysisunmatched
    • Gitunmatched
    • Instrumentationunmatched
    • Machine Learningunmatched
    • Mathematicsunmatched
    • Metadataunmatched
    • Metricsunmatched
    • Object Oriented Programming (OOP)unmatched
    • Parallel Computingunmatched
    • Performance Analysisunmatched
    • Performance Modelingunmatched
    • Process Improvementunmatched
    • Python Programming/Scripting Languageunmatched
    • Software Architectureunmatched
    • Statisticsunmatched
    • Team Playerunmatched
    • Wearablesunmatched
    • Writing Skillsunmatched

    Description

    The Health Sensing Machine Learning Interpretability & Analytics (MLIA) team ensures clinical rigor and contextual trust are at the foundation of Apple's health sensing features. We are looking for an exceptional ML Engineer to help us build the next generation of scalable evaluation infrastructure and lead rigorous investigations into model performance. You will develop cutting-edge tools, synthetic data pipelines, and automated frameworks that ensure our health features are mathematically sound, demographically equitable, and clinically safe. If you are passionate about AI safety, robust software architecture, and pushing the boundaries of ML innovation, come join us! In this role, you will architect and build large-scale evaluation frameworks to interrogate unimodal ML systems and multi-modal foundation models. Beyond infrastructure, you will lead deep-dive ML evaluations, performing failure analysis to uncover performance gaps, reasoning flaws, and edge cases. You will translate findings into actionable insights and work directly with algorithm teams to improve the safety and reliability of our health features. Your work will empower teams across Apple to rapidly evaluate multi-modal sensor fusion while upholding Apples privacy standards.Design robust methodologies and scalable frameworks to assess the performance, reliability, and safety of both traditional ML and foundation models (e.g., LLMs, diffusion models). Drive failure analysis along with building instrumentation to detect clinical hallucinations, reasoning flaws, and edge cases. Expand LLM/diffusion-based data generation pipelines that enable model training and evaluation without exposing real user data. Build data adaptors and visualizers to fuse asynchronous time-series signals (wearables, camera, behavioral metadata). Develop generalizable tools and metrics to discover biases and measure demographic equity across diverse populations Translate evaluation results into actionable engineering insights for GenAI researchers, algorithm leads, and clinical experts.BS in Computer Science, Machine Learning, Statistics, or related field 3+ years of experience in ML Engineering or Applied ML Strong experience in evaluating supervised, unsupervised, LLMs and deep learning models. Proficiency in Python with the ability to write production-grade code (OOP, CI/CD, Git) Hands-on experience in failure analysis, evaluating LLMs and driving subsequent model improvements Experience building data pipelines, inference frameworks, and automated evaluation systems Strong communication skills to articulate complex technical concepts across technical and non-technical audiencesMS/PhD in Computer Science, Machine Learning, Statistics, or related field Experience evaluating LLMs or agentic systems (e.g., LLM-as-a-judge, RAG evaluation) Experience with synthetic data generation and prompt engineering Experience in parallel data processing (Spark, Kubernetes, Airflow) or privacy-preserving ML (Federated Learning) Background in AI Safety, model interpretability, or adversarial testing Interest in digital health and clinical rigor

    Numbers & Facts

    LocationCupertino, CA
    IndustryComputer/IT Services
    Company Size10,000 employees or more
    Year Founded1976
    Websitehttps://www.apple.com/jobs

    About Company

    We bring amazing people together to make amazing things happen.

    We’re a diverse collection of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. The people who work here have reinvented entire industries with the Mac, iPhone, iPad, and Apple Watch, as well as with services, including iTunes, the App Store, Apple Music, and Apple Pay. And the same passion for innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it.

    About Apple

    There’s a place here for every kind of brilliant. Everyone here is an innovator, or an innovator-to-be, no matter what your team or your role. So bring your passion, courage, and original thinking and get ready to share it, because every new product, service, or feature we invent is the result of people working together to make each others’ ideas stronger. Innovation at this level depends on people who represent the variety of the human experience and inspire us with their own fresh perspectives. Together, we’ll do amazing work that can make a difference in people’s lives. Including your own. Learn more about working at Apple.

    Similar Jobs