BS or MS in Computer Science / related fields or equivalent work experience 5+ years in a Site Reliability Engineering or infrastructure-focused engineering role Proficient in Java; proficiency in Go (golang) or Python is also valuable Hands-on experience designing or operating workflow / orchestration / job-scheduling systems, or building substantial automation for stateful, long-running, failure-prone operations Support of internet-facing production services and distributed systems via deployments, on-call, and incident management Experience running large-scale infrastructure with a heavy reliance on automation tooling Real operational experience managing services at scale on Kubernetes Excellent troubleshooting and performance deep-dive analysis Self-motivated and inquisitive, with an aptitude to learn new technologies quickly and effectivelyDirect experience with workflow systems (or comparable durable-execution engines such as Temporal or Cadence): writing workflows and activities, reasoning about retries, signals, queries, heartbeating, and execution history Experience operating or building tooling for distributed databases / storage systems (Cassandra, Redis/Valkey, Kafka, Solr, ZooKeeper, etcd) Understanding of database concepts: consistency models, isolation levels, crash and recovery semantics Operational experience deploying in and running on datacenter and cloud architectures: networking topologies, host placement strategies, and failure modes; design of multi-datacenter systems; failure domains; and wide-area networking Performance engineering (design concepts, profile-guided optimization) Familiarity with mTLS / certificate-based service auth and Kubernetes-native worker deployment Experience replacing legacy or home-grown orchestration with modern, observable platforms Comfort working in a distributed team across multiple sites and time zones. The operations that keep this fleet healthy (rolling upgrades, host replacements, cluster expansions, backup and restore, release qualification) are long-running, multi-step, failure-prone procedures that today rely on internal workflow systems and manual runbooks.