Administer research applications and infrastructure across Google Cloud Platform, AWS, and scientific SaaS environments, including IAM, service accounts, networking, storage, compute, Cloud Run services, containers, registries, DNS, TLS certificates, load balancing, and Identity-Aware Proxy while maintaining separation across development, test, validation, and production environments. Maintain logs, metrics, traces, alerts, dashboards, and automated health checks covering availability, performance, errors, API limits, model and tool failures, token consumption, and cloud cost; create runbooks and support incident triage, escalation, root-cause analysis, backup, restoration, and disaster recovery.