We are looking for a strong SRE resource to work from Rhode Island office 3 days a week (Hybrid). Share strong profile aligning to the JD listed below.
Location : Rhode Island
Required Qualifications
8+ years of Senior Software engineering experience in SRE, DevOps, platform engineering, or related production-systems roles in distributed systems at production scale with active on-call responsibility
Demonstrated experience as an on-call Incident Commander (IC) for P1 or P2 incidents structured leadership updates, not just participant involvement
Experience tuning and validating time-series anomaly detection models in a production observability context this is a Required qualification, not a preferred one; anomaly-based detection is a core function of this role
Strong programming proficiency in Python , React , and Java at production quality capable of writing operational tooling that other engineers will rely on
Hands-on experience designing SLIs, SLOs, and managing error budgets for customer-facing or business-critical services
Deep observability platform experience: Prometheus , Grafana , OpenTelemetry , and at least two of the log aggregation solution (Loki, Splunk, Elasticsearch)
Fleet-scale deployment awareness: familiarity with progressive rollout strategies, blast radius management, and configuration drift as a reliability risk in large unattended node deployments
Strong cloud platform expertise in Google Cloud Platform (GCP) and Rancher K3s.
Advanced Kubernetes operational experience: debugging, resource management, networking policies, and workload failure modes. Experience with AI-assisted tooling and development.
Experience diagnosing and resolving workflow orchestration issues, batch processing failures, scheduler performance problems, and building observability on data pipeline : Apache Airflow and Tidal .
Preferred Qualifications
Experience owning Production Readiness Reviews or service launch gates.
Strong proficiency in transforming large-scale operational and telemetry data into actionable business insights using SQL-based analytics, and reporting frameworks: Google BigQuery, PostgreSQL.
Hands-on chaos or fault injection experience.
TIC (Technical Incident Commander) certification or equivalent structured incident command training
Experience operating distributed systems in retail, pharmacy, healthcare, or other operationally sensitive environments where failures have direct patient or customer impact
LLM integration for operational use cases (alert summarization, runbook suggestion, incident triage assistance) design or implementation experience
Experience with streaming data platforms: Kafka .
Experience with service mesh and traffic management: Istio , Envoy .
Infrastructure-as-code proficiency at production scale: Terraform or Ansible
Position Summary :
Own and drive the end-to-end reliability, availability, and performance of critical retail and pharmacy technology platforms across hybrid cloud and on-premises environments.
Establish and maintain SLI/SLO health, alerting strategies, observability standards, and business-aligned monitoring for the assigned application domain.
Lead production incident response as Incident Commander, drive root cause analysis, postmortems, and continuous reliability improvements.
Partner with engineering, product, and operations teams to embed reliability, resiliency, scalability, and operational readiness into system design and delivery.
Build and optimize automation, self-service capabilities, and operational tooling to eliminate toil, improve efficiency, and reduce manual intervention.
Design and execute proactive reliability initiatives, including production readiness reviews, dependency risk assessments, fault injection, and chaos engineering exercises.
Mentor engineers, champion SRE best practices, and enable teams to independently detect, respond to, and learn from production issues with minimal SRE involvement.
Influence organizational adoption of SLO-driven engineering, observability, incident management, and reliability practices through collaboration, credibility, and measurable outcomes.
Since 1996, RJT has provided successful SAP, Oracle, and IT consulting solutions and staffing services to clients around the world. The new Apolis brings you the same personalized service fortified with a greater array of IT solutions, global expertise, and cost-management strategies.
We are a global IT consultancy that seamlessly integrates experts and leading-edge solutions into your organization so you can focus on what really matters.
Skills
Ansibleunmatched
Apacheunmatched
Artificial Intelligence (AI)unmatched
Automationunmatched
Best Practicesunmatched
Budget Managementunmatched
Business Servicesunmatched
Cloud Computingunmatched
Computer Programmingunmatched
Customer Relationsunmatched
Data Managementunmatched
Debugging Skillsunmatched
DevOpsunmatched
Distributed Computingunmatched
Elasticsearchunmatched
Engineeringunmatched
GCP (Good Clinical Practices)unmatched
Healthcareunmatched
Hybrid Cloudunmatched
Identify Issuesunmatched
Incident Managementunmatched
Incident Responseunmatched
Injectionsunmatched
Integrated Circuits (ICs)unmatched
Leadershipunmatched
Machine Toolunmatched
Mentoringunmatched
On Callunmatched
Pharmacyunmatched
PostgreSQLunmatched
Problem Solving Skillsunmatched
Product/Service Launchunmatched
Production Systemsunmatched
Python Programming/Scripting Languageunmatched
React.jsunmatched
Reliability Engineeringunmatched
Resource Managementunmatched
Retailunmatched
Riskunmatched
Risk Analysisunmatched
Root Cause Analysisunmatched
SQL (Structured Query Language)unmatched
Schedule Developmentunmatched
Software Engineeringunmatched
Splunkunmatched
Standards Strategyunmatched
Telemetryunmatched
Use Casesunmatched
Vehicle Fleetsunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.