3+ years of experience in site reliability or production engineering for large-scale systems-defining and owning SLIs, SLOs, and SLAs; error budgets; incident command and on-call; building and operating production observability (metrics, tracing, logging-e.g., OpenTelemetry, Prometheus, Grafana, Datadog, Dynatrace, Amazon CloudWatch, Azure Monitor, Google Cloud Operations, SolarWinds, Splunk); environment integrity and drift prevention across pre-production and production; and segregation-of-duties controls (least-privilege/RBAC, deploy approvals, secrets management) in partnership with security and risk. Advanced Technical Proficiency: Possess expertise in site reliability and modern production engineering-cloud platform ownership, observability (metrics, tracing, logging), performance and capacity engineering, chaos engineering, and cloud/AI cost engineering-together with applied AI fluency to operate AI and agentic workloads reliably, including AI and Agentic SSDLC, delivering production operations with full automation from discovery to production to operations and all quality checks through the SSDLC lifecycle.