SRE Leader - GPS & YAVA Platform Reliability

Info Way Solutions LLC
  • Greater Toronto Area, CT
  • Instant Apply
2 days ago

Job Description

SRE Leader - GPS & YAVA Platform Reliability

Ld Director, Site Reliability Engineering

PRIMARY OUTCOME
Prevent incidents

PLATFORM SCOPE
GPS and YAVA

ROLE TYPE
Full-time leadership

Role Summary

The SRE Leader will build and lead the Site Reliability Engineering function for the GPS Salesforce CRM and YAVA -Self Service platforms. The primary mandate is to reduce P1/P2 incidents by establishing the people, practices, governance, and engineering capabilities required to identify risk early, control change-related failures, improve detection, and prevent recurring incidents.

What this role owns
The reliability operating model and the SRE team. This leader does not replace application, infrastructure, security, telephony, middleware, cloud, or vendor owners. The role creates shared standards, visibility, accountability, and preventive controls across those teams.

Core Responsibilities

1. Build the SRE Function

  • Define the SRE vision, charter, operating model, engagement model, and implementation roadmap.

  • Recruit, onboard, manage, coach, and retain SRE engineers with the right mix of software, infrastructure, observability, and production engineering expertise.

  • Define SRE competencies, career paths, team objectives, coverage expectations, and engineering standards.

  • Establish clear boundaries between SRE, application development, production support, infrastructure operations, incident management, and vendor teams.

2. Establish Reliability Standards and Governance

  • Define SLIs, SLOs, error budgets, and reliability policies for critical GPS and YAVA journeys.

  • Prioritize incident-oriented indicators such as GPS availability, member-search success, CTI integration success, YAVA transaction success, dependency error rates, and p95/p99 latency.

  • Establish production-readiness reviews and minimum requirements for monitoring, runbooks, capacity, rollback, and ownership.

  • Create a regular reliability review forum and an executive scorecard focused on material risks, incident trends, and corrective-action progress.

3. Lead the Incident Reduction Program

  • Analyze P1/P2 incidents across applications and shared dependencies to identify systemic failure patterns.

  • Create and maintain a prioritized reliability roadmap and backlog based on severity, recurrence, customer or agent impact, and preventability.

  • Ensure every P1/P2 has a defensible root cause, contributing factors, preventive actions, owners, and due dates.

  • Escalate overdue or ineffective preventive actions and drive elimination of repeat failure modes.

4. Strengthen Change and Upgrade Controls

  • Establish risk-based review and validation requirements for changes affecting GPS, YAVA, telephony, DNS, F5, Imperva, MQ, mainframe/DB2, API gateways, cloud, operating systems, databases, and vendors.

  • Require impact assessments, regression evidence, pre-change checks, post-change verification, and tested rollback for high-risk changes.

  • Partner with Change Management to identify changes requiring SRE review or sign-off.

  • Track change failure rate and use incident evidence to improve change controls without creating unnecessary bureaucracy.

5. Drive Observability and Preventive Engineering

  • Set the strategy for end-to-end telemetry, synthetic monitoring, dependency health, alert quality, and service-health correlation.

  • Ensure monitoring reflects actual GPS agent and YAVA customer experiences, not only component status.

  • Sponsor automation for configuration drift, certificates, DNS, routing, load balancer membership, queue health, and post-change validation.

  • Drive investment in resilience patterns, capacity testing, failover validation, and removal of single points of failure.

6. Lead Cross-Team Reliability Accountability

  • Establish working agreements with application, network, security, telephony, middleware, database, mainframe, cloud, and vendor teams.

  • Maintain clear dependency ownership, support contacts, escalation paths, and vendor communication expectations.

  • Use data to surface unresolved risks and secure decisions or investment from senior leadership.

  • Represent platform reliability in major change, readiness, incident, and operational risk forums.

Success Measures

Measure

Expected Direction

P1/P2 incidents

Sustained reduction in count and business impact

Change-induced incidents

Lower percentage and severity of failures caused by changes

Repeat incidents

Reduction by dependency and failure category

Early detection

Higher percentage detected before customer or agent impact

MTTD and MTTR

Improvement in detection and restoration time

Preventive actions

Higher on-time completion and demonstrated effectiveness

Change validation

Higher coverage for in-scope high-risk changes

SLO performance

Improved attainment and disciplined error-budget use

Required Qualifications

  • Demonstrated experience building or materially scaling an SRE function from the ground up.

  • Proven experience directly managing, coaching, and developing SRE engineers.

  • Leadership experience across SRE, production engineering, platform engineering, DevOps, or large-scale technology operations.

  • Strong command of SLIs, SLOs, error budgets, observability, incident management, resilience engineering, and production-readiness practices.

  • Experience driving reliability across organizational boundaries and influencing teams not in the direct reporting line.

  • Strong understanding of distributed systems, networking, DNS, load balancing, TLS, API gateways, cloud platforms, middleware, databases, and vendor dependencies.

  • Executive communication skills with the ability to translate technical risk into customer, agent, operational, and financial impact.

Preferred Experience

  • Enterprise CRM and Salesforce ecosystems.

  • Contact-center technology, CTI, telephony, real-time voice, or conversational AI platforms.

  • IBM MQ, DB2, mainframe integrations, Imperva, F5, Five9, Azure, or Google Cloud.

  • Regulated healthcare or another high-availability enterprise environment.

Numbers & Facts

LocationGreater Toronto Area, CT

Skills

  • Analysis Skillsunmatched
  • Application Programming Interface (API)unmatched
  • Artificial Intelligence (AI)unmatched
  • Automationunmatched
  • Budgetingunmatched
  • Change Controlunmatched
  • Change Managementunmatched
  • Cloud Computingunmatched
  • Coachingunmatched
  • Communication Skillsunmatched
  • Computer Telephony Integration (CTI)unmatched
  • Corrective Actionunmatched
  • Customer Experienceunmatched
  • Customer Relationship Management (CRM)unmatched
  • DNS (Domain Name System)unmatched
  • Database Middleware Softwareunmatched
  • DevOpsunmatched
  • Distributed Computingunmatched
  • Doctor of Nursing Science (DNS)unmatched
  • Establish Prioritiesunmatched
  • F5 Network Softwareunmatched
  • Failoverunmatched
  • GPS (Global Positioning System)unmatched
  • Healthcareunmatched
  • High Availabilityunmatched
  • IBM DB2unmatched
  • IBM WebSphere MQ (Message Queue)unmatched
  • Incident Managementunmatched
  • Leadershipunmatched
  • Load Balancingunmatched
  • Mainframe Computerunmatched
  • Messaging Middlewareunmatched
  • Microsoft Windows Azureunmatched
  • Middlewareunmatched
  • Operating Systemsunmatched
  • Operations Managementunmatched
  • Production Supportunmatched
  • Reliability Engineeringunmatched
  • Requirements Validation/Verificationunmatched
  • Riskunmatched
  • Risk Analysisunmatched
  • Risk Managementunmatched
  • SSL-TLS (Secure Socket Layer - Transport Layer Security)unmatched
  • Salesforce.comunmatched
  • Scorecardingunmatched
  • Software Developmentunmatched
  • Standards Developmentunmatched
  • Strategic Planningunmatched
  • Team Lead/Managerunmatched
  • Technical Operationsunmatched
  • Telemetryunmatched
  • Telephonyunmatched
  • Test Patternsunmatched
  • Time Managementunmatched
  • Validation Testingunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder