Lead Devops Engineer

Way2B1
  • San Francisco, California
    3 days ago

    Job Description

    Our Company and the Role

    This is a hands-on Lead DevOps Engineer role responsible for designing, operating, and evolving a highly available, multi-tenant platform on AWS. You will work closely with software engineering to deploy, operate, and scale production systems while driving improvements in reliability, automation, and performance.

    You will lead a small DevOps team of 1–2 engineers while remaining primarily hands-on. You will serve as the team's technical lead, set standards and priorities, mentor and develop engineers, and take ownership of the most difficult infrastructure and production problems.

    You will also help introduce and operationalize AI/LLM capabilities within the platform.

    This role reports to the Chief Technology Officer (CTO) and is part of the engineering leadership team.

    What You'll Be Responsible For

    • Design, build, and operate scalable, highly available infrastructure in AWS
    • Own and evolve infrastructure as code (Terraform) across all environments
    • Operate and optimize Aurora PostgreSQL, including replication, failover, and performance tuning
    • Operate ECS (Fargate), ECR, and containerized services
    • Operate Kafka-based event streaming systems
    • Manage Auto Scaling Groups and EC2-based workloads
    • Design and maintain CI/CD pipelines using Buildkite
    • Build automation to eliminate manual operational work
    • Manage and secure secrets and access using Vault, AWS Secrets Manager, and IAM
    • Partner with engineering teams to improve system reliability and performance
    • Lead, mentor, and manage a small DevOps team
    • Set technical direction, standards, and priorities for the DevOps function
    • Drive cost optimization across AWS infrastructure
    • Operate systems behind Cloudflare, including WAF, CDN, and traffic management

    Production Reliability & Incident Ownership

    • Own production incident response end-to-end, including triage, mitigation, and coordination
    • Lead high-severity outage response under pressure
    • Serve as an escalation point for complex production issues
    • Drive root cause analysis (RCA) and enforce follow-up actions
    • Continuously improve system resilience and recovery mechanisms

    Observability & System Insight

    • Design and operate end-to-end observability across metrics, logs, and tracing
    • Build high-signal monitoring, alerting, and dashboards
    • Define and enforce SLIs, SLOs, and alerting standards
    • Reduce alert fatigue and improve the signal-to-noise ratio

    AI / LLM Systems — Emerging Area

    • Experience using AI/agentic developer tools, such as Claude Code, Cursor, or similar tools, to accelerate DevOps workflows and improve engineering efficiency
    • You don't have to be an expert with AI yet, but you must have the desire to learn quickly and become proficient
    • This is an area we're heavily investing in as a company

    What You Bring

    • Deep experience operating production systems on AWS, including ECS/Fargate, EC2, networking, and IAM
    • Expert-level Terraform experience managing infrastructure at scale
    • Strong experience with containerized applications and distributed systems, such as Kafka
    • Experience operating multi-tenant, highly available systems
    • Proven ownership of production on-call and experience resolving critical incidents
    • Strong systems fundamentals, including Linux, networking, and debugging
    • Strong scripting ability in Bash, Python, or an equivalent language
    • Experience designing and operating CI/CD systems
    • Strong understanding of security best practices, including IAM and secrets management
    • Demonstrated technical leadership and experience mentoring engineers
    • Experience managing engineers, or clear readiness and interest in taking on direct people management
    • Strong communication and prioritization skills

    Nice to Have

    • Experience operating multi-region or globally distributed systems
    • Experience working with Cloudflare at scale
    • Experience optimizing high-throughput or event-driven systems
    • Experience leading a small infrastructure, platform, SRE, or DevOps team
    • Experience operating or integrating LLM/AI services in production environments, including tracing and evaluation with OpenTelemetry, LangSmith, Langfuse, or equivalent tools
    • Experience managing the performance, cost, and reliability of LLM workloads, including latency, token usage, rate limiting, and fallbacks

    Oh Yeah, and We Can Offer You

    • A tight-knit team of motivated, dedicated individuals who work together without ego
    • Extensive access to and engagement with leadership
    • An Agile working methodology
    • Competitive salary
    • Access to new technologies
    • 401(k)
    • Cell phone reimbursement
    • Health, dental, and vision benefits

    Numbers & Facts

    LocationSan Francisco, California

    Skills

    • Agile Programming Methodologiesunmatched
    • Amazon Elastic Compute Cloud (EC2)unmatched
    • Amazon Web Services (AWS)unmatched
    • Apache Kafkaunmatched
    • Artificial Intelligence (AI)unmatched
    • Automationunmatched
    • Autoscalingunmatched
    • Bash Scriptingunmatched
    • Best Practicesunmatched
    • Cellular Telephoneunmatched
    • Communication Skillsunmatched
    • Content Delivery Network (CDN)unmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Improvementunmatched
    • Continuous Integrationunmatched
    • Cost Controlunmatched
    • DevOpsunmatched
    • Distributed Applicationsunmatched
    • Distributed Computingunmatched
    • Engineering Managementunmatched
    • Establish Prioritiesunmatched
    • Failoverunmatched
    • High Availabilityunmatched
    • High Throughputunmatched
    • Incident Responseunmatched
    • Leadershipunmatched
    • Linux Operating Systemunmatched
    • Mentoringunmatched
    • Metricsunmatched
    • Network Debuggingunmatched
    • On Callunmatched
    • People Managementunmatched
    • Performance Managementunmatched
    • Performance Tuning/Optimizationunmatched
    • PostgreSQLunmatched
    • Process Improvementunmatched
    • Production Systemsunmatched
    • Programming Toolsunmatched
    • Python Programming/Scripting Languageunmatched
    • Reimbursementunmatched
    • Reliability Engineeringunmatched
    • Replication and Remote Mirroringunmatched
    • Reporting Dashboardsunmatched
    • Root Cause Analysisunmatched
    • Scalable System Developmentunmatched
    • Signal-to-noise Ratio (SNR)unmatched
    • Software Engineeringunmatched
    • System Operationsunmatched
    • Systems Engineeringunmatched
    • Systems Reliabilityunmatched
    • Team Lead/Managerunmatched
    • Technical Leadershipunmatched
    • Traffic Shapingunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder