Apples Services Engineering organization (ASE) is seeking experienced site reliability engineers to build and maintain the workflow orchestration and automation platform for our Data Services fleet. Data Services SRE operates Cassandra, Redis/Valkey, Kafka, Solr, and our Coordination systems (ZooKeeper, etcd, Parallax) across Apples data centers worldwide. These systems form the platform upon which iCloud and many other internet services at Apple are built. The operations that keep this fleet healthy (rolling upgrades, host replacements, cluster expansions, backup and restore, release qualification) are long-running, multi-step, failure-prone procedures that today rely on internal workflow systems and manual runbooks. You will work to replace these legacy systems with durable, observable, resumable workflows. Your work will reduce on-call toil, cut incident recovery time, and turn tribal operational knowledge into self-documenting code. In ASE, your work benefits hundreds of millions of users and is critical to the reliability of some of the most visible current and future Apple features. The ASE Data Services workflow effort builds automation that is safe, reliable, observable, and resumable. This work requires an innovative spirit and an extraordinary degree of care and rigor in engineering. These workflows operate on production databases at massive scale, where a mishandled rolling upgrade or host replacement has real customer impact. You will support the workflow engine layer underlying Apples most critical database systems which power all of Apples internet services. You will design and build workers and workflow definitions, composing operations across existing execution layers: host and pod provisioning, cluster topology discovery, and our monitoring and alerting systems. You will migrate operational logic off home-grown workflow tooling onto a durable execution model with retry, signal, query, and full execution history. This role requires excellent communication, the ability to partner closely with developers, SRE, and platform teams, and a high degree of customer focus when engaging with the internal stakeholders who will run these workflows every day. As a distributed team, the ability to work effectively with colleagues based in other locations is essential; experience in this area is a plus. Prior experience building or operating workflow/orchestration systems, or operating distributed databases and storage systems at scale, is recommended.Design, build, and operate durable workflow automation for Data Services operations (rolling upgrades, host replacement, cluster expansion and shrink, backup and restore, and release qualification). Write workflow and activity code that composes existing systems (per-node management APIs, provisioning APIs, topology discovery, monitoring/alerting, notifications) into safe, resumable operations. Apply core SRE practice (monitoring, alerting, and incident management) to the automation platform itself and to the operations it drives. Migrate operational logic off legacy workflow systems and manual runbooks, triaging what ports directly, what should be rewritten, and what can be retired. Build the developer and operator experience around these workflows: CLIs, execution visibility, and the patterns other SREs use to submit, monitor, signal, and debug running operations. Understand database operational concepts deeply enough to automate them safely: drain and recovery semantics, ring balancing, streaming, health validation, and the failure modes of multi-step operations against a live cluster. Operate across bare-metal, virtualized (EC2), and containerized (Kubernetes) platforms, deploying workers to Kube and reasoning about cross-datacenter connectivity, host placement, and failure domains. Partner with platform teams across the org to keep our work aligned, portable, and contributing back where it makes sense.BS or MS in Computer Science / related fields or equivalent work experience 5+ years in a Site Reliability Engineering or infrastructure-focused engineering role Proficient in Java; proficiency in Go (golang) or Python is also valuable Hands-on experience designing or operating workflow / orchestration / job-scheduling systems, or building substantial automation for stateful, long-running, failure-prone operations Support of internet-facing production services and distributed systems via deployments, on-call, and incident management Experience running large-scale infrastructure with a heavy reliance on automation tooling Real operational experience managing services at scale on Kubernetes Excellent troubleshooting and performance deep-dive analysis Self-motivated and inquisitive, with an aptitude to learn new technologies quickly and effectivelyDirect experience with workflow systems (or comparable durable-execution engines such as Temporal or Cadence): writing workflows and activities, reasoning about retries, signals, queries, heartbeating, and execution history Experience operating or building tooling for distributed databases / storage systems (Cassandra, Redis/Valkey, Kafka, Solr, ZooKeeper, etcd) Understanding of database concepts: consistency models, isolation levels, crash and recovery semantics Operational experience deploying in and running on datacenter and cloud architectures: networking topologies, host placement strategies, and failure modes; design of multi-datacenter systems; failure domains; and wide-area networking Performance engineering (design concepts, profile-guided optimization) Familiarity with mTLS / certificate-based service auth and Kubernetes-native worker deployment Experience replacing legacy or home-grown orchestration with modern, observable platforms Comfort working in a distributed team across multiple sites and time zones
| Location | Seattle, WA |
| Industry | Computer/IT Services |
| Company Size | 10,000 employees or more |
| Year Founded | 1976 |
| Website | https://www.apple.com/jobs |
We’re a diverse collection of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. The people who work here have reinvented entire industries with the Mac, iPhone, iPad, and Apple Watch, as well as with services, including iTunes, the App Store, Apple Music, and Apple Pay. And the same passion for innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it.
There’s a place here for every kind of brilliant. Everyone here is an innovator, or an innovator-to-be, no matter what your team or your role. So bring your passion, courage, and original thinking and get ready to share it, because every new product, service, or feature we invent is the result of people working together to make each others’ ideas stronger. Innovation at this level depends on people who represent the variety of the human experience and inspire us with their own fresh perspectives. Together, we’ll do amazing work that can make a difference in people’s lives. Including your own. Learn more about working at Apple.
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.
Free resume templatesImprove your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.
Free resume builder