Strengthen observability across application, data, orchestration, and infrastructure layers by improving metrics, logs, traces, dashboards, alerts, and service-level indicators to track availability, error rates, and incident response time while collaborating with site reliability engineering teams to improve incident response, on-call readiness, runbooks, and troubleshooting workflows. Ability to learn and apply new languages, tools, and frameworks as the problem requires, such as Ruby, Go, or TypeScript and Vue in adjacent parts of the stack, along with excellent written communication and asynchronous collaboration skills demonstrated through thoughtful code review, mentoring, incident communication, and context sharing.