Working closely with the Availability Command Center (ACC), Tier 1 Operations, Application Engineering teams, Infrastructure teams, and external vendors, this position provides elevated level of technical expertise for application incidents, operational issues, and service stability initiatives. Advanced working knowledge of Kubernetes, OpenShift, Docker containers, and microservice-based architectures, including the ability to troubleshoot complex distributed-system failures, evaluate platform health, recommend recovery strategies, and influence operational best practices.