Lead Cloud Engineering and Production Operations Engineer

A10 Networks

San Jose, CA

JOB DETAILS
SALARY
$110,000–$155,000 Per Year
SKILLS
Amazon Web Services (AWS), Ansible, Applications Security, Architectural Services, Automation, Autoscaling, Bash Scripting, Best Practices, Capacity Management, Chef (Configuration Management), Cloud Architecture, Cloud Computing, Communication Skills, Computer Networks, Computer Science, Configuration Management, Continuous Deployment/Delivery, Continuous Integration, Cost Control, Cross-Functional, Customer Relations, DevOps, Docker, Documentation, Federal Contracts, High Availability, Hybrid Cloud, Identify Issues, Incident Response, Information Technology & Information Systems, Jenkins, Leadership, Mentoring, Microsoft Windows Azure, Network Security, Performance Metrics, Production Costing, Production Support, Production Systems, Puppet (Configuration Management), Python Programming/Scripting Language, Reliability Engineering, Root Cause Analysis, Scalable System Development, Scripting (Scripting Languages), Software Engineering, Software as a Service (SaaS), Standard Operating Procedures (SOP), Team Player, Technical Leadership, Testing, Time Management, Windows PowerShell
LOCATION
San Jose, CA
POSTED
Today

Lead Cloud Engineering and Production Operations Engineer
This role acts as a hands-on technical lead, driving cloud engineering initiatives, automating infrastructure, and ensuring high-availability and performance across customer-facing systems. The Lead Engineer will collaborate with IT, DevOps, and Software Engineering teams to build secure, scalable environments that support continuous delivery and rapid innovation.

Reporting to the Associate Director of IT and Infrastructure, this position combines deep technical execution with mentoring responsibilities-balancing architectural vision with day-to-day operational excellence.

Key Responsibilities:

Cloud Infrastructure and Engineering

  • Design, deploy, and manage hybrid and cloud infrastructures (OCI, AWS, Azure, on-prem) to support production and enterprise systems
  • Implement infrastructure-as-code (IaC) using Terraform or CloudFormation to ensure repeatable, secure, and automated deployments
  • Develop and maintain CI/CD-ready environments that support rapid build, test, and release cycles for engineering teams
  • Partner with network and security teams to implement resilient, compliant architectures
Production Operations and Reliability
  • Serve as technical lead for production systems, ensuring stability, performance, and scalability
  • Establish monitoring, logging, and alerting frameworks to improve visibility and reduce mean time to detection (MTTD) and resolution (MTTR)
  • Participate in incident response, root cause analysis, and reliability improvement efforts
  • Collaborate with Engineering and SRE teams to define SLIs, SLOs, and performance metrics for critical services
Automation and CI/CD Enablement
  • Develop and enhance deployment pipelines (e.g., Jenkins, GitLab, ArgoCD) to automate software delivery and environment provisioning
  • Embed security, compliance, and testing gates into CI/CD workflows
  • Implement configuration management and orchestration tools such as Ansible, Chef, or Puppet to manage infrastructure at scale
  • Drive efficiency through self-healing systems, auto-scaling, and infrastructure automation
Operational Leadership and Collaboration
  • Lead day-to-day production operations activities, mentoring junior engineers on cloud and reliability best practices
  • Act as a technical bridge between Infrastructure, Security, and Application Engineering teams
  • Contribute to capacity planning, cost optimization, and production readiness reviews
  • Maintain documentation, runbooks, and standard operating procedures for production systems
Qualifications:
  • Bachelor's degree in Computer Science, Information Systems, or equivalent experience
  • 7+ years of experience in cloud and infrastructure engineering, with at least 2-3 years in a lead or senior engineer capacity
  • Deep expertise in OCI (preferred) AWS or Azure (networking, compute, storage, IAM, and monitoring)
  • Proven experience with production-scale operations and hybrid cloud deployments
  • Proficiency in:
    • Infrastructure-as-code (Terraform, CloudFormation)
    • CI/CD and DevOps pipelines (Jenkins, GitLab, ArgoCD)
    • Containers and orchestration (Kubernetes, Docker)
    • Observability tools (Datadog, Prometheus, Grafana, ELK)
    • Scripting languages (Python, Bash, PowerShell)
  • Strong troubleshooting skills and the ability to lead through high-impact incidents
  • Excellent communication and collaboration skills across cross-functional teams


Preferred Experience:
  • Experience supporting high-availability SaaS or production environments
  • Knowledge of FinOps, cloud governance, and cost optimization practices
  • Familiarity with DevSecOps principles, Zero Trust, and automated compliance frameworks
  • Exposure to AI/ML pipeline infrastructure or high-throughput data systems


AI Use Guidelines for Interviews: Our interviews are designed to reflect your own skills and thinking. The use of AI or recording tools during live interviews is not permitted unless explicitly invited by the interviewer or approved in advance as part of a reasonable accommodation. If these tools are used inappropriately or in a way that misrepresents your work, your application may not move forward in the process.

Why Join Us:

This is a hands-on leadership role for an engineer who thrives at the intersection of cloud architecture, automation, and reliability. As the Lead Cloud Engineering and Production Operations Engineer, you'll have direct impact on how the company delivers, scales, and secures its core technology platforms-enabling speed, stability, and innovation across the enterprise.

A10 Networks is an equal opportunity employer and a VEVRAA federal subcontractor. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability status, protected veteran status, or any other characteristic protected by law. A10 also complies with all applicable state and local laws governing nondiscrimination in employment.

#LI-AN1 - Hybrid

Targeted compensation guideline: $110,000 - $155,000. Compensation will vary based on number of factors, including market demand for specific skills, role type, job level, and individual qualifications. Final salary offers are determined by considerations including, but not limited to, subject matter expertise, demonstrated skill level, relevant experience, geographic location, education, certifications, and training.

About the Company

A

A10 Networks