Our Client, an IT Services and Consultant company, is looking for a Site Reliability Engineer (SRE) - AI & Automation for their Dallas, TX / Scottsdale, AZ / Hybrid location.
Responsibilities:
Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.
Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.
Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.
Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments.
Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics to meet reliability and performance objectives.
Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement.
Requirements:
Site Reliability Engineering (SRE) - Reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.
Kubernetes Platform Engineering - 5+ years of Strong hands-on experience with GKE and Rancher RKE2, multi-cluster management, troubleshooting, and performance optimization.
Cloud & Infrastructure Automation - Strong experience in GCP, Terraform, Helm, GitHub, CI/CD, and production-grade automation.
Software Development - 5+ years of Advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows).
Observability & Monitoring - Splunk, Grafana, Datadog, AppDynamics, alerting, and platform health monitoring.
API & Microservices Engineering - Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.
AI-Driven Operations (AIOps) - Applying LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident triage, automation, and operational workflows.
Years of Experience: 14.00 Years of Experience
Why Should You Apply?
Health Benefits
Referral Program
Excellent growth and advancement opportunities
Numbers & Facts
Location
Dallas, TX
Skills
Analysis Skillsunmatched
Application Programming Interface (API)unmatched
Artificial Intelligence (AI)unmatched
Automationunmatched
Best Practicesunmatched
Cloud Computingunmatched
Computer Programmingunmatched
Computer Securityunmatched
Continuous Deployment/Deliveryunmatched
Continuous Improvementunmatched
Continuous Integrationunmatched
Cross-Functionalunmatched
Disaster Recoveryunmatched
Failoverunmatched
GCP (Good Clinical Practices)unmatched
GitHubunmatched
GraphQLunmatched
Health Planunmatched
High Availabilityunmatched
High Reliabilityunmatched
Identify Issuesunmatched
Incident Managementunmatched
Incident Responseunmatched
Information Technology Consultingunmatched
Javaunmatched
Microservicesunmatched
Network Operations Centerunmatched
Node.jsunmatched
Performance Tuning/Optimizationunmatched
Python Programming/Scripting Languageunmatched
REST (Representational State Transfer)unmatched
Reliability Engineeringunmatched
Software Developmentunmatched
Splunkunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.