Job Description
Job Title: DevOps & Site Reliability Lead
Location: Deerfield IL, 60015 - Hybrid Role 3 days
Employment Type : Full Time
Experience: 10 years
Must Have Technical/Functional Skills
- Cloud & Platform Engineering (Expert Level)
- Deep expertise in Microsoft Azure, including:
- Compute (VMs, App Services, Azure Container Apps)
- Containers & Orchestration (AKS, Docker)
- Networking (VNETs, Private Endpoints, Application Gateway, Load Balancers)
- Storage, Azure Key Vault, Azure Monitor, Log Analytics
- Proven experience designing enterprise grade, highly available cloud platforms
- Strong understanding of hybrid and multi cloud architectures (AWS / GCP exposure preferred)
DevOps & Engineering Excellence
- Advanced experience with Azure DevOps and CI/CD pipeline architecture
- Infrastructure automation using Terraform (modules, state management, governance)
- Strong scripting skills (PowerShell, Bash)
- GitOps concepts, branching strategies, release orchestration
- Site Reliability Engineering (Leadership Level)
- Ownership of platform reliability, resiliency, and performance
Definition and governance of:
- SLIs, SLOs, SLAs
- Error budgets and reliability metrics
- Advanced observability strategy:
- Metrics, logs, traces, alerts, dashboards using Dynatrace
- Incident response leadership, RCA facilitation, and long term remediation planning
- Experience operating 99.9% 99.99% availability systems
Containers, APIs & Integration
- Leadership-level experience with AKS-based platforms, ingress, and scaling strategies
- Understanding of microservices, API-led and event-driven architectures
- Familiarity with Azure Integration Services (Service Bus, Event Hub, API Management)
Security, Compliance & Cost
- Secure cloud design using Key Vault, managed identities, RBAC
- Cost optimization (FinOps mindset) across cloud infrastructure
Roles & Responsibilities
- Act as Lead SRE for client's Digital platforms, owning reliability and stability outcomes
- Define and enforce SRE standards, best practices, and operating models
- Architect and govern highly available, scalable cloud platforms
- Lead the design and implementation of CI/CD and IaC strategies
- Establish proactive monitoring, alerting, and incident prevention mechanisms
- Own major incident leadership, RCA execution, and corrective action tracking
- Partner with application, security, and architecture teams to build reliability by design
- Drive automation to reduce toil and improve operational efficiency
- Mentor and coach SRE and DevOps engineers across teams
- Influence roadmap decisions with a reliability, scalability, and cost lens
Skills
Amazon Web Services (AWS)unmatched
Application Programming Interface (API)unmatched
Applications Securityunmatched
Automationunmatched
Bash Scriptingunmatched
Best Practicesunmatched
Budgetingunmatched
Cloud Architectureunmatched
Cloud Computingunmatched
Coachingunmatched
Continuous Deployment/Deliveryunmatched
Continuous Integrationunmatched
Corrective Actionunmatched
Cost Controlunmatched
DevOpsunmatched
GCP (Good Clinical Practices)unmatched
High Availabilityunmatched
Hybrid Cloudunmatched
Identity Data Managementunmatched
Incident Responseunmatched
Leadershipunmatched
Mentoringunmatched
Metricsunmatched
Microservicesunmatched
Microsoft Windows Azureunmatched
Operational Improvementunmatched
Operational Strategyunmatched
Reliability Engineeringunmatched
Reporting Dashboardsunmatched
Scripting (Scripting Languages)unmatched
Security Architectureunmatched
Service Level Agreement (SLA)unmatched
Software Engineeringunmatched
VMS Operating Systemunmatched
Virtual Machine (VM)unmatched
Windows PowerShellunmatched
Level up your application
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.
Free resume templatesImprove your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.
Free resume builder