Apple Inc logo

Software Engineer, Infrastructure Services (Cloud Network Fault Tolerance)

Apple Inc
  • California, CA
    3 days ago

    Job Description

    Do you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple's Information Systems and Technology (IS&T) organization. IS&T is the engine behind everything Apple does for customers and for the people who build for them. It's Apple's central nervous system. Supporting 2.5 billion active Apple devices, processing billions of secure transactions, and keeping the technology that defines modern life running flawlessly, IS&T makes the impossible feel effortless."

    Do you love building solutions to handle global complexity and immense scale? Imagine what you could do here. Infrastructure Services is part of IS&T and the foundation of Apples global network operations - managing data center equipment and systems to deliver compute, storage, and networking services for teams across Apple, including its internal developer community. From individual facilities to a worldwide network, Infrastructure Services ensures the technology underneath everything works without question.

    This is not a traditional network engineering role. This is a software engineering role for builders who love deep technical challenges, have strong fundamentals in distributed systems and fault-tolerant design, and want to turn cutting-edge ideas into production systems running at massive scale. We welcome early-career engineers with exceptional technical depth, if you have the intellectual horsepower and hunger to solve problems that dont have textbook answers yet, we want to talk to you. You will join a team chartered to make Apples hyper-scale cloud network dramatically more fault-tolerant by building systems that detect and remediate failures automatically, with minimal or no human intervention. This means designing and building the sensing, decision-making, and remediation layers that allow the network to identify anomalies, localize root cause, and take corrective action in real time.

    Youll work on problems like: how do you detect a degrading network path before it causes an outage? How do you safely automate remediation actions on live production infrastructure without introducing new risk? How do you build a system that learns from past incidents to prevent recurrence? These are open, hard problems, and you will be expected to bring rigorous thinking, creativity, and strong engineering execution to solve them.

    This role is ideal for someone who loves taking a deep, formal understanding of distributed systems, control theory, graph algorithms, or fault-tolerant design patterns and turning it into resilient, real-world software running in production at enormous scale. You should be comfortable reading research papers and translating relevant ideas into working systems, while also being pragmatic about the constraints of operating in a live, hyper-scale environment.

    If you have exceptional technical depth from your academic background, competitive programming, research experience, or personal projects, and youre hungry to work on problems most companies arent yet equipped to solve, this team is built for you.Design and build automated detection, diagnosis, and remediation systems for network faults across a hyperscale cloud environment Develop the decision logic and safety mechanisms that allow automated systems to take corrective action on production infrastructure with confidence Analyze failure patterns and systemic weaknesses across the network and design software solutions to eliminate or mitigate them Build simulation, chaos testing, and fault-injection tooling to proactively surface weaknesses before they cause real incidents Apply formal reasoning about distributed systems, failure modes, and control systems to build robust, production-grade software Translate emerging research and industry innovation in resilience engineering into practical systems that work at Apples scale Bring new ideas and approaches to the team, this is a group expected to push the boundaries of whats possible in network resiliency Prototype and experiment rapidly, validating ideas against real production data and failure scenarios Partner with Cloud Network production and reliability teams to understand real-world failure modes and operational constraints Work with peer infrastructure teams to ensure self-healing systems integrate safely and effectively across the broader platformBachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience 4 to 6+ years of professional software engineering experience designing, building, and operating production-grade distributed systems and backend infrastructure Deep foundation in computer science fundamentals, including distributed systems architecture, concurrency models, graph algorithms, and systems design Strong proficiency in at least one systems-level or high-performance language, such as Go, C++, Rust, or Python Direct experience designing and building fault-tolerant mechanisms, including automated self-healing, active remediation, circuit breaking, load shedding, and blast-radius mitigation for network services Demonstrated ability to model complex failure domains, handle network partitions and split-brain scenarios, and reason rigorously about system behavior under extreme load and degradation Practical experience with chaos engineering, fault injection, simulation-based testing, and stress testing in production or staging environments Track record of technical ownership, including authoring design documents, driving code reviews, and leading post-mortem root cause analysesMaster's or Ph.D. in Computer Science, Distributed Systems, Networking, or a related technical discipline Deep domain knowledge in cloud networking architectures, Software-Defined Networking (SDN) control planes, L3/L4 routing protocols, overlay networks, and Linux networking constructs (such as eBPF, XDP, or OVS) Hands-on experience implementing or tuning distributed consensus algorithms (such as Raft or Paxos) and closed-loop control systems or reconciliation controllers Experience architecting high-throughput telemetry pipelines and automated anomaly detection systems for real-time network health analysis Proven ability to digest academic papers and industry research, translating state-of-the-art resilience concepts into production systems History of notable open-source contributions, technical publications, or patent filings in distributed networking and systems reliability

    Numbers & Facts

    LocationCalifornia, CA
    IndustryComputer/IT Services
    Company Size10,000 employees or more
    Year Founded1976
    Websitehttps://www.apple.com/jobs

    About Company

    We bring amazing people together to make amazing things happen.

    We’re a diverse collection of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. The people who work here have reinvented entire industries with the Mac, iPhone, iPad, and Apple Watch, as well as with services, including iTunes, the App Store, Apple Music, and Apple Pay. And the same passion for innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it.

    About Apple

    There’s a place here for every kind of brilliant. Everyone here is an innovator, or an innovator-to-be, no matter what your team or your role. So bring your passion, courage, and original thinking and get ready to share it, because every new product, service, or feature we invent is the result of people working together to make each others’ ideas stronger. Innovation at this level depends on people who represent the variety of the human experience and inspire us with their own fresh perspectives. Together, we’ll do amazing work that can make a difference in people’s lives. Including your own. Learn more about working at Apple.

    Skills

    • Academic Backgroundunmatched
    • Algorithmsunmatched
    • Analysis Skillsunmatched
    • Appleunmatched
    • Building Systemsunmatched
    • C++ Programming Languageunmatched
    • Cloud Architectureunmatched
    • Cloud Computingunmatched
    • Code Reviewsunmatched
    • Competitive Researchunmatched
    • Computer Engineeringunmatched
    • Computer Scienceunmatched
    • Concurrencyunmatched
    • Control Systemsunmatched
    • Corrective Actionunmatched
    • Data Managementunmatched
    • Design Documentunmatched
    • Design Patterns Programming Methodologiesunmatched
    • Distributed Computingunmatched
    • Distributed Control Systems (DCS)unmatched
    • Electrical Engineeringunmatched
    • Engineeringunmatched
    • Failure Analysisunmatched
    • Graph Theoryunmatched
    • High Throughputunmatched
    • Information Technology & Information Systemsunmatched
    • Infrastructure as a Service (IaaS)unmatched
    • Injectionsunmatched
    • Linux Networksunmatched
    • Machine Toolunmatched
    • Mail Processingunmatched
    • Network Administration/Managementunmatched
    • Network Architect Softwareunmatched
    • Network Architecture/Engineeringunmatched
    • Network Designunmatched
    • Network Operations Centerunmatched
    • Network Performance/Analysisunmatched
    • Network Protocolsunmatched
    • Network Softwareunmatched
    • Network Systemsunmatched
    • Open Sourceunmatched
    • Patent Applicationsunmatched
    • Pattern Analysisunmatched
    • Problem Solving Skillsunmatched
    • Production Systemsunmatched
    • Prototypingunmatched
    • Python Programming/Scripting Languageunmatched
    • Reconciliationunmatched
    • Riskunmatched
    • Routing Protocolsunmatched
    • Rust Programming Languageunmatched
    • Simulationunmatched
    • Software Designunmatched
    • Software Engineeringunmatched
    • Stress Testingunmatched
    • System Architectureunmatched
    • System Integration (SI)unmatched
    • Systems Reliabilityunmatched
    • Team Playerunmatched
    • Technical Publicationsunmatched
    • Telemetryunmatched
    • Testingunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder