Site Reliability Engineer

Expert In Recruitment Solutions
  • malvern, PA
  • Quick Apply
30+ days ago

Job Description

Site Reliability Engineer
hybird - malvern, pa
needs at least 8 years experience within the US

Job Description
The Site Reliability Engineer (SRE) is responsible for improving the reliability, resiliency, observability, and operational excellence of Client's Cash & Money Movement ecosystem. This role serves as the reliability leader for the department, partnering with product and engineering teams to identify gaps, prevent incidents, improve recovery, and ensure critical client journeys remain highly available and resilient.

Key Responsibilities
Observability & Monitoring
  • Own and maintain a single-pane-of-glass dashboard for application, platform, dependency, and client journey health.
  • Improve SLOs, SLIs, alerts, dashboards, and monitoring standards.
  • Ensure proactive detection of client-impacting issues using logs, metrics, traces, and synthetic monitoring.
Reliability & Incident Management
  • Improve MTTD, MTTR, and overall service reliability.
  • Maintain incident response playbooks and alerting standards.
  • Facilitate blameless postmortems, root cause analysis, and track corrective actions through closure.
  • Analyze trends and recurring failure patterns to prevent repeat incidents.
Resilience Engineering
  • Lead FMEA assessments for critical applications and journeys.
  • Identify single points of failure and partner with teams on remediation plans.
  • Conduct Game Days, chaos testing, failover testing, and recovery exercises.
  • Validate multi-region, multi-AZ, and disaster recovery capabilities.
Safe Change & Operational Excellence
  • Define reliability standards and operational guardrails.
  • Review production readiness of high-risk changes.
  • Drive adoption of safe deployment practices such as canary releases, feature flags, and automated rollback mechanisms.
Community of Practice & Reliability Leadership
  • Build and lead the Cash & Money Movement SRE Community of Practice.
  • Drive engagement, knowledge sharing, and reliability culture across the organization.
  • Identify and mentor application-level SRE champions/POCs.
  • Facilitate weekly reliability forums, office hours, and operational reviews.
  • Educate teams on SRE best practices, observability, incident management, resilience testing, and safe change principles.
  • Partner closely with Danlin Hibay's SRE and operational excellence organizations to stay aligned with enterprise standards, emerging tools, lessons learned, and engineering best practices.
  • Act as the liaison between Cash & Money Movement and enterprise SRE communities to bring recommendations, standards, and innovations back to product teams
Key Deliverables
  • Unified Cash & Money Movement Reliability Dashboard
  • Journey Health Dashboard (Add Bank, Transfers, Wires, ACH, Direct Deposit, Cash Plus, etc.)
  • SLO/SLI Framework and Alert Standards
  • FMEA Library and Resiliency Test Plans
  • Incident Playbooks and Postmortem Reviews
  • Reliability Community of Practice
  • Reliability Maturity Assessments and Executive Reporting
Success Measures
  • Reduced Sev 1/2/3 incidents
  • Reduced MTTD and MTTR
  • 100% critical applications with SLOs, dashboards, and actionable alerts
  • Completion of FMEA and resiliency testing for critical journeys
  • Timely closure of postmortem action items
  • Improved reliability, availability, and client experience across Cash & Money Movement.
  • Active and engaged reliability community across Cash & Money Movement
Operating Model
This is a Hub-and-Spoke SRE model, SRE defines what "good" looks like and drives continuous improvement while engineering teams remain accountable for execution and results.

SRE owns
  • Reliability standards and best practices
  • Observability and dashboards
  • Assessments, FMEA, and resilience testing
  • Incident reviews and postmortems
  • Community of Practice
  • Education, coaching, and governance
Product Teams own
  • Reliability backlog execution
  • Remediation and implementation
  • Operational outcomes
  • Service health and reliability improvements

Numbers & Facts

Locationmalvern, PA

Skills

  • Best Practicesunmatched
  • Community of Practice (CoP)unmatched
  • Continuous Improvementunmatched
  • Corrective Actionunmatched
  • Customer Experienceunmatched
  • Disaster Recoveryunmatched
  • Ecosystemsunmatched
  • Failoverunmatched
  • Failure Mode and Effects Analysis (FMEA)unmatched
  • High Availabilityunmatched
  • Incident Managementunmatched
  • Incident Responseunmatched
  • Leadershipunmatched
  • Mentoringunmatched
  • Metricsunmatched
  • Operational Auditunmatched
  • Process Improvementunmatched
  • Product Engineeringunmatched
  • Reliability Analysisunmatched
  • Reliability Engineeringunmatched
  • Reporting Dashboardsunmatched
  • Riskunmatched
  • Root Cause Analysisunmatched
  • Standards Developmentunmatched
  • Test Plan/Scheduleunmatched
  • Testingunmatched
  • Time Managementunmatched
  • Trend Analysisunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder