What you'll do
1Drive the rollout, integration, localization, and operational support of DevOps and high-availability platforms across international regions, ensuring platform capabilities are delivered reliably and on schedule.
2Evaluate infrastructure differences across international public cloud, private cloud, and data center environments, including compute, networking, storage, Kubernetes, security, and compliance requirements. Design and implement practical adaptation plans.
3Support the international adoption of internal DevOps platforms, including service release systems, change management platforms, and large-scale operations tooling. Improve CI/CD pipelines, infrastructure automation, and release operations workflows.
4Implement high-availability capabilities for international environments, including incident response playbooks, failover mechanisms, disaster recovery validation, and chaos engineering practices. Track remediation actions and continuously improve service reliability.
5Participate in production operations and incident response for international environments. Quickly identify and troubleshoot issues related to releases, infrastructure, and platform integration, and drive root-cause analysis and closed-loop improvements.
6Maintain deployment guides, operational runbooks, incident response playbooks, and standard operating procedures to improve delivery quality and operational efficiency across international regions.
7Work closely with central platform R&D teams, as well as international infrastructure, networking, security, and business teams. Manage project progress, dependencies, and risks, and provide feedback to improve platform capabilities for international scenarios.
Qualifications
1Bachelor’s degree or above in Computer Science, Software Engineering, or a related field.
2Hands-on experience in DevOps, SRE, cloud platforms, or infrastructure engineering, with strong delivery and troubleshooting capabilities.
3Solid understanding of Linux, computer networking, and common infrastructure components. Able to independently perform deployment, configuration, and issue diagnosis.
4Familiar with cloud-native and Infrastructure-as-Code technologies, such as Kubernetes, Helm, and Terraform. Experience with KubeVela is a plus.
5Proficient in at least one programming or scripting language, such as Go, Java, Python, or Shell. Able to build automation tools and read/debug backend service code.
6Familiar with at least one major public cloud or enterprise private cloud environment. Understand concepts such as multi-region deployment, network connectivity, access control, and security isolation.
7Familiar with monitoring, logging, and alerting systems. Experience with Prometheus, Grafana, ELK/OpenSearch, OpenTelemetry, or similar tools is preferred.
8Fluent in both English and Chinese, with the ability to communicate effectively in a cross-region, international team environment.
9Strong ownership, execution, and communication skills. Able to translate central platform designs into stable, maintainable implementations in local international environments.
| Location | Palo Alto, California |
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.
Free resume templatesImprove your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.
Free resume builder