We have positions for SRE to be filled with our portfolio and it is a hybrid position, preferably at Shea, AZ immediately.
Interview would be conducted in multiple rounds with different set of panels focused on Technical Skills & Experience. Kindly request to perform early screenings on your end, to avoid delay and valuable time from the panel.
Site Reliability Engineer (SRE) Job Requirements
We are seeking a Site Reliability Engineer with strong experience in cloud-native operations, observability, automation, and production support for large-scale enterprise applications.
Skillset required:
3-5 years of experience in Site Reliability Engineering, Production Operations, or Platform Engineering supporting large-scale, high-performance applications across hybrid environments (on-premises and cloud).
3-5 years of experience developing automation scripts and building Application Performance Management (APM) dashboards to monitor end-to-end transaction journeys.
Hands-on programming experience (2+ years) with one or more languages such as Go, Python, Java, or Rust.
Working knowledge of relational and NoSQL databases including Oracle, SQL Server, PostgreSQL, MongoDB, Redis, ClickHouse, PL/SQL, or time-series databases.
Experience with cloud migration and containerization initiatives using GCP, AWS, Azure, Rancher, OpenShift, or similar platforms.
Experience managing containerized applications in Kubernetes environments such as GKE, RKE, or AKS.
Strong experience implementing observability solutions using Open Telemetry (OTEL), distributed tracing, monitoring, and incident management.
Familiarity with GraphQL frameworks such as Apollo, Prisma, or Hasura.
Strong networking fundamentals including TCP/IP, HTTP, DNS, load balancing, and service mesh technologies.
Experience participating in 24x7 on-call rotations and meeting incident response SLAs.
Preferred Qualifications:
Experience managing highly available, customer-facing platforms with a focus on reliability, automation, and operational excellence.
Hands-on experience with monitoring and observability tools such as Splunk, Dynatrace, AppDynamics, Grafana, and Prometheus.
Experience with CI/CD and Agile tools such as Rally, Confluence, and related DevOps platforms.
Knowledge of in-memory caching technologies, especially Redis.
Strong troubleshooting and debugging skills across distributed systems and API gateway architectures.
Experience with Google Cloud services including GCS, Cloud SQL, Spanner, and BigQuery.