Google Cloud Site Reliability Engineer (GCP SRE)

ICONMA, LLC

Buffalo Grove, IL

JOB DETAILS
SALARY
$15–$46.42 Per Hour
SKILLS
Administrative Management, Agile Modeling, Analysis Skills, Budgeting, Cloud Computing, Cloud Storage, Communication Skills, Data Management, DevOps, GCP (Good Clinical Practices), GitHub, Health Plan, High Availability, Incident Management, Information Technology Consulting, Machine Learning, Metrics, Microsoft Visual Studio, On Call, Presentation/Verbal Skills, Problem Solving Skills, Process Modeling, Product Lifecycle, Production Support, Production Systems, Python Programming/Scripting Language, Reliability Engineering, Reporting Dashboards, Root Cause Analysis, SQL (Structured Query Language), Service Level Agreement (SLA), ServiceNow, Software Engineering, Splunk, System Operations, Systems Engineering, Tableau, Team Player, Time Management
LOCATION
Buffalo Grove, IL
POSTED
14 days ago
Our client, a IT Services and Consulting company, is looking for a Google Cloud Site Reliability Engineer (GCP SRE)”: for their Buffalo Grove, IL/Hybrid location.
 
Responsibilities:
  • Responsible for Incident Detection & Logging and meeting agreed SLA for incident tickets.
  • Responsible for Bridge Activation & Communication (P1–P2).  
  • Postmortem Preparation (Within 24–72 Hours) & Root Cause Analysis.
  • Responsible for critical monitoring activities ,Problem Management & Grafana Integration.
  • Participate in oncall rotations, handle incidents, and drive timely mitigation and recovery.
  • Automating operational work so services can scale without manual toil also operating highly available, low latency & secure systems.
  • Defining and measuring reliability through SLIs/SLOs and error budgets.
  • Build and maintain observability: metrics, logs, traces, dashboards, and alerts for critical services.
  • Tune alerting to reduce noise while ensuring rapid detection of user impacting issues.
  • Lead or contribute to post incident reviews and root cause analysis and ensure follow up actions are implemented to prevent recurrence.
  • Added Advantage if resource is familiar on Tools Tidal, Service Now, Xmatters, Abinitio, Tableau, Opsgenie & Zeke.
 
Requirements:
  • 8+ years in Site Reliability Engineering, DevOps, or Cloud Engineering
  • Strong experience supporting production GCP environments
  • Experience with enterprise incident management and cloud operations
  • Knowledge/experience in GCP (Big Query, Cloud storage, Dataproc, GKE, Airflow/Composer ,Pub-sub, Cloud function, Cloud SQL etc).
  • Knowledge/experience in Github & Visual Studio code.
  • Knowledge/experience in MS-Copilot.
  • Knowledge/experience in Prometheus, Grafana & Splunk.
  • Knowledge in Python/Pyspark/Machine learning is an added advantage
  • Clear written and verbal communication, particularly under pressure (e.g., during incidents).
  • Ability to collaborate across multiple teams and influence engineering practices through expertise rather than authority.
  • Strong communication, analytical ,Knowledge on entire Incident management life cycle process, Agile model experience, and problem-solving skills
  • EDP-SRE (Google Site Reliability Engineer)
  • Site Reliability Engineers combined software engineering with systems and infrastructure operations to build and run large, reliable, scalable services.
  • Years of Experience: 10.00 Years of Experience
Skills:
  • Category                                 Name               Required         Importance     Experience
  • Data Mgt_Administration         GCP                 Yes                  1
  • DevOps                                   SRE                 Yes                  1
  • EPS_NEW                               GITHUB          Yes                  1
 
Why Should You Apply?

About the Company

I

ICONMA, LLC