Senior Production Engineer

Anduril Industries

Costa Mesa, CA

JOB DETAILS
SALARY
$191,000–$287,000 Per Year
SKILLS
Circuit Breakers, Cloud Computing, Computer Security, Customer Support/Service, Debugging Skills, Distributed Computing, Environmental Compliance, Identify Issues, Production Support, Production Systems, Replication and Remote Mirroring
LOCATION
Costa Mesa, CA
POSTED
Today

ABOUT THE TEAMThe SRE team owns reliability and infrastructure for Anduril's cloud deployments. We operate Kubernetes clusters, Terraform infrastructure, and observability platforms across 10+ production environments supporting active defense contracts. When platform services break under real operational load, we're the team that fixes them — often at the code level, not just the config level.ABOUT THE JOBWe are looking for a Senior Production Engineer to join our team in Costa Mesa, CA (or DC) . In this role, you will be responsible for diagnosing and fixing stability vulnerabilities in core platform services that cause cascading failures in multi-tenant cloud deployments. You will write production Go to implement resilience patterns — leader election, circuit breakers, failure domain isolation — directly in service code. This will require deep experience with distributed systems, debugging complex failure modes across service boundaries, and writing production-quality Go. If you are someone who thrives on fixing hard reliability problems in live systems rather than building greenfield, this role is for you.WHAT YOU'LL DODiagnose and fix stability vulnerabilities in core platform services that cause cascading failures under multi-replica, multi-tenant operationImplement resilience patterns (leader election, circuit breakers, failure domain isolation) directly in service codeDesign multi-replica support for services that currently assume single-instance operationCollaborate with service owners on contract testing and upgrade validationTrace cascading failures across service boundaries and drive them to root-cause fixesContribute to observability platform improvements to support service stabilityLight infrastructure work: Terraform/Kubernetes changes to support service fixes (~20% of time)REQUIRED QUALIFICATIONSProduction-quality Go — you'll be modifying core platform services, not writing scriptsPractical experience with distributed systems: leader election, consensus, replication, failure modesKubernetes — enough to understand how services run (not necessarily cluster administration)Debugging complex systems — tracing cascading failures across service boundaries4+ years in SRE, platform engineering, or backend development rolesMust be a U.S. Person due to required access to U.S. export controlled information or facilitiesNICE-TO-HAVE QUALIFICATIONSRust (some platform services use it)Experience fixing reliability problems in production services (not just building greenfield)Familiarity with gRPC service architecturesHashiCorp Consul or similar service discovery/meshFedRAMP/IL5 compliance environment experienceArgoCD / GitOps workflowsUS Salary Range: $191,000 – $287,000 USDBenefitsAt Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you're supported in health, recovery, and whatever comes next.#J-18808-Ljbffr

About the Company

A

Anduril Industries