Responsibilities:Designs solutions to visualize key production support metrics enabling Operational Readiness and Site Reliability Engineer teams to identify scenarios requiring interventionDevelops software solutions and/or improved processes to address work identified as 'toil' by collaborating with key partners to identify, track and remediate processes to free time allocated to reliabilityPartners with Development and Infrastructure teams to create error budget policies prioritizing reliability stories that fall below Service Level Objective (SLO) thresholds and suggests code optimizations, additional instrumentation and/or logging structures to gain service reliability visibilityIdentifies and plans for capacity bottlenecks, vulnerabilities and opportunities for reliability improvement, such as low level error rates and 'noise', and reduces manual support effort and/or improves system reliabilityAssesses monitoring for new changes with development partners and works with monitoring tools team to monitor dashboards and enhance application and system monitoring designsEngages as a subject matter expert in incident triage efforts, failure scenario modelling and works with the Problem Manager to diagnose root causes for complex/high impact incident/problem management investigationsCollaborates with Development and Infrastructure teams to understand technical solutions and develop Service Level Indicators and SLOs to measure/improve the reliability of the services they supportChampions modern SRE, observability, and AIOps practices by driving automation, reliability engineering, and operational excellence across platform services. Demonstrated experience transforming traditional infrastructure operations, systems administration, or platform support functions into modern engineering-led operating models utilizing Site Reliability Engineering (SRE), observability, automation, and AIOps practices.·Proven experience implementing and operationalizing SRE principles, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, reliability metrics, incident management modernization, and continuous service improvement programs.