Job title: DevOps Engineer Kubernetes & Cloud Platform
Location: San Jose, CA 95101
Experience: 5 8 Years
Education: Bachelor's/Master's in Engineering
Job Summary
We are seeking an experienced DevOps Engineer / Kubernetes & Cloud Platform Engineer to design, automate, and support highly scalable cloud-native infrastructure. The ideal candidate will have strong hands-on expertise with Kubernetes, AWS, Azure, Terraform, Python automation, CI/CD, GitOps, and cloud observability.
The candidate will be responsible for building and maintaining reliable Kubernetes platforms, automating infrastructure and deployments, implementing monitoring solutions, and improving platform performance, security, and operational efficiency.
Key Responsibilities
- Design, deploy, administer, and optimize Kubernetes environments, including AWS EKS and Azure AKS.
- Develop and maintain Infrastructure as Code (IaC) using Terraform.
- Build Python-based automation for infrastructure provisioning, operational tasks, and self-healing processes.
- Implement and maintain CI/CD pipelines using Jenkins, GitLab, Harness, and ArgoCD.
- Drive GitOps-based deployment and platform engineering practices.
- Manage Docker-based containerized workloads and Kubernetes deployments.
- Implement Kubernetes scaling and optimization strategies using HPA and Karpenter.
- Configure and support Istio Service Mesh for secure and reliable microservice communication.
- Develop enterprise monitoring and observability solutions using Prometheus, Grafana, and Datadog.
- Troubleshoot production infrastructure, Kubernetes workloads, networking, deployments, and performance issues.
- Implement cloud security best practices across AWS and Azure environments.
- Automate infrastructure lifecycle management and reduce manual operational effort.
- Support highly available, scalable, and secure cloud-native platforms.
- Collaborate with development, security, and infrastructure teams to modernize CI/CD and cloud operations.
- Identify opportunities for performance optimization, reliability improvements, and cloud cost reduction.
Required Skills
- 5 8 years of experience in DevOps, Cloud Engineering, SRE, or Platform Engineering.
- Strong hands-on experience with Kubernetes, preferably EKS and/or AKS.
- Strong experience with AWS and Azure cloud platforms.
- Hands-on expertise with Terraform and Infrastructure as Code.
- Strong Python scripting/automation experience.
- Experience with Docker and Kubernetes administration.
- Experience with ArgoCD and GitOps.
- Strong CI/CD experience with Jenkins, GitLab, or Harness.
- Experience with Prometheus, Grafana, and Datadog.
- Knowledge of Istio Service Mesh.
- Experience with Kubernetes scaling technologies such as HPA and Karpenter.
- Strong Linux administration and troubleshooting skills.
- Understanding of cloud security, networking, IAM/RBAC, and highly available architectures.
Preferred Qualifications
- Experience with enterprise Platform Engineering and SRE practices.
- Experience implementing cloud-native security and observability.
- Knowledge of automated remediation and reliability engineering.
- Experience supporting regulated environments such as Banking, Mortgage, or Insurance.
- Bachelor's or Master's degree in Engineering or a related technical field.