We are seeking an experienced Site Reliability Engineer (SRE) to support large-scale, high-performance applications running in a hybrid environment (on-premises and cloud). The ideal candidate will have strong experience in cloud infrastructure, Kubernetes, observability, automation, and production operations.
Requirements
Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud).
Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys.
Experience working with Programming languages such as Go, Python, Java, Rust etc.
Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any time-series databases
Experience in transitioning platforms to the cloud and Containerization – GCPand Rancher
Experience maintaining containerized app in GKE/RKE/AKE environments.
Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution.
Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc...).
Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.Service reliability/operation experience running large-scale, high-performance applications in a hybrid environment (on-prem and cloud).
Experience in writing automation scripts and building dashboards for Application Performance management to manage Transaction journeys.
Experience working with Programming languages such as Go, Python, Java, Rust etc.
Working knowledge on with one or more databases- Oracle, SQL Server, Redis, Clickhouse, postgres, Mongo or any time-series databases
Experience in transitioning platforms to the cloud and Containerization – GCPand Rancher
Experience maintaining containerized app in GKE/RKE/AKE environments.
Experience Implementing Cloud observability using OTEL to enable real-time monitoring, distributed tracing and incident resolution.
Experience working with specific GraphQL Framework (Apollo, Prisma, Hasura etc...).
Experience using knowledge of networking protocols such as TCP/IP, HTTP, DNS, Load balancing and service mesh to troubleshoot issues in high pressure situations.
Numbers & Facts
Location
Scottsdale, Arizona
Skills
Application Performance Managementunmatched
Automationunmatched
Cloud Computingunmatched
DNS (Domain Name System)unmatched
GraphQLunmatched
HTTP (HyperText Transport Protocol)unmatched
Home Automationunmatched
Identify Issuesunmatched
Javaunmatched
Load Balancingunmatched
Microsoft SQL Serverunmatched
Network Protocolsunmatched
Oracle Databaseunmatched
PostgreSQLunmatched
Programming Languagesunmatched
Python Programming/Scripting Languageunmatched
Redisunmatched
Reliability Engineeringunmatched
Reporting Dashboardsunmatched
Rust Programming Languageunmatched
Scripting (Scripting Languages)unmatched
Software Administrationunmatched
TCP/IP (Transmission Control Protocol/Internet Protocol)unmatched
Time Trackingunmatched
Transaction Processing/Managementunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.