Our client, a IT Services and Consulting company, is looking for a Google Cloud Site Reliability Engineer (GCP SRE)": for their Buffalo Grove, IL/Hybrid location.
Responsibilities:
Responsible for Incident Detection & Logging and meeting agreed SLA for incident tickets.
Responsible for Bridge Activation & Communication (P1-P2).
Postmortem Preparation (Within 24-72 Hours) & Root Cause Analysis.
Responsible for critical monitoring activities ,Problem Management & Grafana Integration.
Participate in oncall rotations, handle incidents, and drive timely mitigation and recovery.
Automating operational work so services can scale without manual toil also operating highly available, low latency & secure systems.
Defining and measuring reliability through SLIs/SLOs and error budgets.
Build and maintain observability: metrics, logs, traces, dashboards, and alerts for critical services.
Tune alerting to reduce noise while ensuring rapid detection of user impacting issues.
Lead or contribute to post incident reviews and root cause analysis and ensure follow up actions are implemented to prevent recurrence.
Added Advantage if resource is familiar on Tools Tidal, Service Now, Xmatters, Abinitio, Tableau, Opsgenie & Zeke.
Requirements:
8+ years in Site Reliability Engineering, DevOps, or Cloud Engineering
Strong experience supporting production GCP environments
Experience with enterprise incident management and cloud operations
Knowledge/experience in Github & Visual Studio code.
Knowledge/experience in MS-Copilot.
Knowledge/experience in Prometheus, Grafana & Splunk.
Knowledge in Python/Pyspark/Machine learning is an added advantage
Clear written and verbal communication, particularly under pressure (e.g., during incidents).
Ability to collaborate across multiple teams and influence engineering practices through expertise rather than authority.
Strong communication, analytical ,Knowledge on entire Incident management life cycle process, Agile model experience, and problem-solving skills
EDP-SRE (Google Site Reliability Engineer)
Site Reliability Engineers combined software engineering with systems and infrastructure operations to build and run large, reliable, scalable services.
Years of Experience: 10.00 Years of Experience
Skills:
Category Name Required Importance Experience
Data Mgt_Administration GCP Yes 1
DevOps SRE Yes 1
EPS_NEW GITHUB Yes 1
Why Should You Apply?
Health Benefits
Referral Program
Excellent growth and advancement opportunities
Numbers & Facts
Location
Buffalo Grove, IL
Skills
Administrative Managementunmatched
Agile Modelingunmatched
Analysis Skillsunmatched
Budgetingunmatched
Cloud Computingunmatched
Cloud Storageunmatched
Communication Skillsunmatched
Data Managementunmatched
DevOpsunmatched
GCP (Good Clinical Practices)unmatched
GitHubunmatched
Health Planunmatched
High Availabilityunmatched
Incident Managementunmatched
Information Technology Consultingunmatched
Machine Learningunmatched
Metricsunmatched
Microsoft Visual Studiounmatched
On Callunmatched
Presentation/Verbal Skillsunmatched
Problem Solving Skillsunmatched
Process Modelingunmatched
Product Lifecycleunmatched
Production Supportunmatched
Production Systemsunmatched
Python Programming/Scripting Languageunmatched
Reliability Engineeringunmatched
Reporting Dashboardsunmatched
Root Cause Analysisunmatched
SQL (Structured Query Language)unmatched
Service Level Agreement (SLA)unmatched
ServiceNowunmatched
Software Engineeringunmatched
Splunkunmatched
System Operationsunmatched
Systems Engineeringunmatched
Tableauunmatched
Team Playerunmatched
Time Managementunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.