Participate in production incident response, perform systematic troubleshooting and root-cause analysis, and drive corrective and preventive actions through blameless post-incident reviews. Support cloud services covering compute, storage, networking, IAM/security, load balancing, monitoring, logging, backup, and disaster recovery.