We are seeking a Senior DevOps Engineer to lead the design, implementation, and optimization of our AWS and GCP infrastructure and development operations. This role will be instrumental in scaling our platform, hardening security, and improving engineering velocity through automation and DevOps best practices. You’ll collaborate closely with engineering, product, and QA teams to ensure that our environments are secure, observable, and high-performing.
What You'll Do
Design, build, and maintain infrastructure across AWS and GCP (EC2/Compute Engine, RDS/Cloud SQL, VPC, S3/Cloud Storage, IAM) using Infrastructure as Code (Terraform)
Maintain and improve CI/CD pipelines in GitHub Actions, including blue/green (AWS CodeDeploy) and in-place deployment strategies, automated database migrations, and release-train gating tied to Linear
Set up and manage monitoring, logging, and alerting across cloud providers (e.g., CloudWatch, GCP Cloud Monitoring/Logging, Datadog)
Collaborate with the security team to define and implement security protocols, incident response plans, and continuous monitoring strategies
Coordinate with security and engineering teams to schedule and execute regular penetration tests, and lead remediation efforts for identified vulnerabilities
Design and operate blue/green deployment infrastructure (AWS CodeDeploy, Auto Scaling Groups/Launch Templates, ALB) across environments, including live production cutovers with zero/minimal downtime
Plan and execute major database version upgrades and zero-downtime cutovers (e.g., Aurora MySQL major-version upgrades) at production scale, including pre-cutover validation, rollback planning, and post-cutover verification
Manage Terraform Cloud–based infrastructure workflows (PR-based plan review and gated apply) across a multi-account AWS organization (separate prod/QA/staging accounts, per-tier IAM roles, per-account ECR)
Implement and manage network-level security controls across cloud providers, including VPC/VPC routing, firewall rules, security groups, and GCP firewall/IAM policies
Deploy and maintain SIEM tools or similar security monitoring systems to centralize logs and alert on anomalous behavior
Harden system access, secret management, and permission models across environments
Optimize performance, scalability, and cost-efficiency of our infrastructure
Support production environments and participate in on-call rotations
Document systems, processes, runbooks, and security best practices
Required Qualifications
5+ years of experience in a DevOps or Site Reliability Engineering (SRE) role
Deep hands-on experience with both AWS and GCP, especially compute, managed databases (RDS/Cloud SQL), VPC/networking, and IAM across providers
Strong proficiency with Terraform (or similar IaC) across multi-cloud environments — managing provider-specific modules/state for AWS and GCP