Apple Inc logo

Sr. / Staff ML Engineer, FM Training Integration - ML Compute

Apple Inc

  • Santa Clara, CA
  • 30+ days ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Analysis Skillsunmatched
    • Appleunmatched
    • Benchmarkingunmatched
    • Best Practicesunmatched
    • Building Systemsunmatched
    • Cloud Computingunmatched
    • Code Reviewsunmatched
    • Computer Scienceunmatched
    • Cross-Functionalunmatched
    • Data Managementunmatched
    • Data Modelingunmatched
    • Debugging Skillsunmatched
    • Deep Learningunmatched
    • Distributed Computingunmatched
    • Dockerunmatched
    • GPU (Graphics Processing Unit)unmatched
    • Hardware Designunmatched
    • JAX (Java API for XML)unmatched
    • Large-Scale Systemsunmatched
    • Machine Learningunmatched
    • Machine Toolunmatched
    • Memory Hardwareunmatched
    • Performance Managementunmatched
    • Performance Modelingunmatched
    • Performance Tuning/Optimizationunmatched
    • Process Improvementunmatched
    • Python Programming/Scripting Languageunmatched
    • Resource Utilizationunmatched
    • Software Engineeringunmatched
    • Systems Scalabilityunmatched

    Description

    We are looking for a ML Engineer to join our ML Compute team to help improve the efficiency, scalability, and reliability of model training and inference workloads in the cloud. In this role, you will lead the integration of large-scale ML workloads with cloud infrastructure, working cross-functionally with ML engineers, infrastructure engineers, and researchers to optimize performance, improve system efficiency, and drive high utilization of accelerator resources. We are a group of engineers to support training foundation models at Apple! We build infrastructure to support training foundation models with general capabilities such as understanding and generation of text, images, speech, videos, and other modalities and apply these models to Apple products. We are looking for engineers who are passionate about building systems that push the frontier of deep learning in terms of scaling, efficiency, and flexibility and delight millions of users in Apple products.Own the integration of large-scale model training workloads with accelerator-based cloud infrastructure, ensuring scalable and reliable execution. Drive performance optimization across the ML stack, including data pipelines, model execution, and distributed systems, to improve throughput, latency, and hardware utilization. Design and run benchmarks to evaluate model performance and infrastructure configurations, using results to guide optimization efforts. Build and improve tooling for observability, profiling, and debugging to increase visibility and reliability of ML workloads. Collaborate cross-functionally with ML engineers, infrastructure engineers, and researchers to improve system efficiency and scalability. Establish and promote best practices for performance tuning and resource utilization. Drive high-quality design and code reviews, share best practices, and elevate engineering standards across the team.5+ years of experience in software engineering, ML infrastructure, or related domains. Hands-on experience with machine learning workflows, including training, evaluation, and inference at scale. Proficiency in Python and experience with at least one major ML framework (e.g., PyTorch or JAX). Experience with cloud-based infrastructure and distributed systems (e.g., containers, orchestration, storage, and networking). Bachelor's degree in Computer Science, Engineering, or a related field.Experience working with accelerator-based systems (e.g., GPUs/TPUs), including performance tuning an debugging of ML workloads. Hands-on experience with distributed training or inference at scale (e.g., data, model, or pipeline parallelism). Experience optimizing large-scale ML systems, including bottleneck analysis across compute, memory, and networking. Familiarity with profiling, tracing, and benchmarking tools for ML workloads (e.g., PyTorch Profiler, NVIDIA Nsight). Experience building or operating ML infrastructure using containerization and orchestration frameworks (e.g., Docker, Kubernetes). Advanced degree in Computer Science, Engineering, or a related field.

    Numbers & Facts

    LocationSanta Clara, CA
    IndustryComputer/IT Services
    Company Size10,000 employees or more
    Year Founded1976
    Websitehttps://www.apple.com/jobs

    About Company

    We bring amazing people together to make amazing things happen.

    We’re a diverse collection of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. The people who work here have reinvented entire industries with the Mac, iPhone, iPad, and Apple Watch, as well as with services, including iTunes, the App Store, Apple Music, and Apple Pay. And the same passion for innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it.

    About Apple

    There’s a place here for every kind of brilliant. Everyone here is an innovator, or an innovator-to-be, no matter what your team or your role. So bring your passion, courage, and original thinking and get ready to share it, because every new product, service, or feature we invent is the result of people working together to make each others’ ideas stronger. Innovation at this level depends on people who represent the variety of the human experience and inspire us with their own fresh perspectives. Together, we’ll do amazing work that can make a difference in people’s lives. Including your own. Learn more about working at Apple.

    Similar Jobs

    See more jobs