Our client, a IT Services and Consulting company, is looking for a Site Reliability Engineer (SRE) for their Plano, TX/ Atlanta, GA/ Middletown, NJ location.
Responsibilities:
Design and execute performance, load, stress, failover, disaster recovery, and end-to-end testing strategies.
Assess application and infrastructure reliability through KPI, SLA, SLO, and availability monitoring.
Conduct capacity planning, failure simulations, and recovery testing to ensure business continuity.
Automate infrastructure provisioning, deployments, monitoring, and remediation processes.
Monitor system health, troubleshoot incidents, and perform root cause analysis (RCA).
Collaborate with development, infrastructure, and security teams to improve platform reliability and performance.
Implement security controls, vulnerability remediation, and compliance best practices.
Requirements:
Seeking an experienced Site Reliability Engineer (SRE) to drive application and infrastructure reliability across cloud environments.
The ideal candidate will have strong expertise in end-to-end testing, performance testing, load testing, failover validation, failure recovery, and operational resilience, with a deep understanding of both infrastructure and application architectures.
Strong experience with Kubernetes / AKS administration and troubleshooting.
Hands-on experience with Azure/AWS/GCP, networking, and cloud infrastructure.
Proficiency in Linux administration and system performance tuning.
Experience with Docker and containerization technologies.
Strong understanding of security, cloud compliance, and vulnerability remediation.
Experience with monitoring and observability tools such as Prometheus, Grafana, Splunk, Dynatrace, or Azure Monitor.
Knowledge of CI/CD and automation practices."
Top Skills:
Kubernetes / AKS administration and troubleshooting.
Hands-on experience with Azure/AWS/GCP, networking, and cloud infrastructure
Proficiency in Linux administration and system performance tuning.