Principal DevOps Engineer

Document Advisor Inc
  • CA
    10 days ago

    Job Description

    Don't miss your chance to join a global team thats transforming the way people travel as a Principal DevOps Engineer. About iVisa At iVisa we believe that traveling should be simple. That's why 2.6M+ travelers have chosen us to facilitate their visas, passports, and other travel documents. We are the easiest, fastest, and simplest solution in the market. Our company is growing 80% year on year. We know our biggest strength is our people and we're looking for the right new team members to help propel our culture and achieve our goals. Above all else, we always have fun! About The Role Own the reliability, security, cost and evolution of iVisas production and staging infrastructure end to end. This is the primary infrastructure role: you are the technical owner of our AWS accounts, Kubernetes clusters, GitOps and CI pipelines, Cloudflare edge, observability stack and internal DevOps tooling; the escalation point for every engineering team; and the manager and mentor of the Site Reliability Engineer who reports to you. Why iVisa Collaborative, friendly, and diverse culture: We foster an inclusive and vibrant atmosphere, featuring a dynamic and international environment with flat hierarchies and exceptionally amiable colleagues. Truly remote-first work environment: work from anywhere or everywhere - we encourage global travel. Access to continuous learning and cutting-edge AI tools, supported by dedicated budgets for professional development. Extended Family Leave policy: Our policy covers all birthing parents, non-birthing parents, and adopting parents. Thrive in a highly tech-savvy company equipped with cutting-edge tools and the power to make a substantial impact. Join us in our commitment to the planet and sustainability: For every iViser, we plant one tree, allowing you to contribute to our environmental initiatives. Rest and Relaxation: We offer flexible PTO for all team members. As a Principal DevOps Engineer at iVisa, you will make an impact by: Platform and cloud Running, patching and upgrading EKS clusters, node groups and add-ons with minimal downtime; planning and communicating maintenance windows. Owning the Terraform estate as the single source of truth: eliminate drift, review and apply infrastructure merge requests. Operating managed data services: RDS MySQL upgrades and parameter tuning, ElastiCache scaling, Redshift users/groups/WLM, S3 lifecycle and access logging, snapshot replication to the backup account. Lead their roadmap items (for example TLS enforcement and auth-plugin migration ahead of MySQL 9). Managing cloud cost: find idle or over-provisioned resources, right-size, retire obsolete distributions, domains and services. Delivery and GitOps Owning GitLab CI/CD and the self-hosted runners; keeping pipelines fast and reliable across all product repositories; maintaining the shared CI images. Owning Argo CD, the Helm chart library and the review-environment lifecycle; unblock developers whose deployments fail. Provisioning infrastructure for new services (buckets, CDNs, IAM, DNS, routing, environment configuration) with the web, AI/ML, data, mobile and growth teams. Reliability and incident response Primary on-call for infrastructure in a rotation shared with the SRE: respond to alerts, lead incidents, restore service, write root-cause analyses and turn them into alerts, runbooks or fixes. Owning the observability stack: metrics, log retention and performance, alert rules, dashboards as code, uptime checks, Sentry organization. Owning disaster-recovery posture: backups, restore testing, multi-replica and multi-node scheduling, capacity for traffic spikes. Security and access Owning the edge: Cloudflare WAF rules, bot, credential-stuffing and enumeration mitigation, IP blocking, Under Attack decisions, Worker deployments. Managing identity and access end to end: onboarding and offboarding for VPN, Vault, Keeper, GitLab and database users; least-privilege IAM; SSO; image vulnerability scanning and policy enforcement. Handling DNS, registrar and certificate operations (Route53, MarkMonitor, BIMI/SPF/DKIM records, wildcard certificates). Platform engineering and enablement Maintaining and extending the in-house Go DevOps tooling and the Runway service catalog. Keeping runbooks and the wiki current (DDoS, VPN, environment configuration, Kubernetes operations); answering requests in #devops-comms; reviewing infrastructure-touching merge requests from other teams. Managing, mentoring and growing the Site Reliability Engineer; set priorities for the DevOps board. Using AI tooling (Claude, MCP servers, agent-friendly repository docs) to automate toil and speed up investigations. What makes you a great fit for this role: 5+ years in DevOps, SRE or platform engineering, including 2+ years as the primary or lead owner of a production Kubernetes platform. Expert Kubernetes on AWS EKS: upgrades, node lifecycle, networking/CNI, ingress (Traefik or similar), RBAC, policy engines, debugging node and dataplane failures. Strong Terraform: modules, multi-environment layouts, state management, drift remediation. Broad AWS: EKS, EC2/Auto Scaling, IAM (including IRSA), VPC, RDS MySQL, ElastiCache, S3, CloudFront, ECR, Route53, ALB. GitOps with Argo CD and Helm chart authoring; GitLab CI with self-hosted Kubernetes runners. Cloudflare in production: DNS, WAF and custom rules, caching, Workers. Observability with Prometheus, Grafana and Loki; alert design; incident command and blameless RCAs. Working proficiency in Go (our internal tooling is Go) and fluent Linux/Bash; comfortable reading PHP, Python and TypeScript. Secrets and identity: HashiCorp Vault or equivalent, rotation, least-privilege design. Networking fundamentals: DNS, TLS, VPN/WireGuard, load balancing, HTTP. Clear written communication: maintenance notices, incident updates and RCAs read by non-infrastructure stakeholders. Available for on-call and occasional off-hours maintenance windows with US Central overlap. Preferred Qualifications External Secrets Operator, Kyverno, Crossplane, Backstage. NATS or a similar event bus; webhook-driven automation. Operating Laravel/PHP-FPM with Horizon and Redis queue workloads. Airflow, OpenMetadata, n8n, Tableau and Redshift data-platform operations. Ansible, Packer, Taskfile; Trivy or Harbor image scanning. FinOps experience; multi-account AWS (we run iVisa and HelloGov accounts). Experience managing or mentoring an engineer. Basic Spanish. iVisa is committed to building a diverse, inclusive, and respectful workplace. We provide equal employment opportunities to all and do not discriminate on the basis of race, gender identity, age, sexual orientation, religion, national origin, disability, or any other protected characteristic.

    Numbers & Facts

    LocationCA

    Skills

    • Amazon CloudFrontunmatched
    • Amazon Elastic Compute Cloud (EC2)unmatched
    • Amazon Simple Storage Service (S3)unmatched
    • Amazon Web Services (AWS)unmatched
    • Ansibleunmatched
    • Artificial Intelligence (AI)unmatched
    • Automationunmatched
    • Autoscalingunmatched
    • Bash Scriptingunmatched
    • Budget Managementunmatched
    • Cachingunmatched
    • Cloud Computingunmatched
    • Communication Skillsunmatched
    • Content Delivery Network (CDN)unmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • Cost Controlunmatched
    • DNS (Domain Name System)unmatched
    • Data Managementunmatched
    • Data Recoveryunmatched
    • Debugging Skillsunmatched
    • Denial of Service (DoS)unmatched
    • DevOpsunmatched
    • Disaster Recoveryunmatched
    • Diversityunmatched
    • Establish Prioritiesunmatched
    • Gardeningunmatched
    • HTTP (HyperText Transport Protocol)unmatched
    • IP (Internet Protocol)unmatched
    • Identity Data Managementunmatched
    • Incident Responseunmatched
    • Laravelunmatched
    • Linux Operating Systemunmatched
    • Load Balancingunmatched
    • MCP - Microsoft Certified Professionalunmatched
    • Machine Toolunmatched
    • Mentoringunmatched
    • Metricsunmatched
    • MySQLunmatched
    • On Callunmatched
    • Onboardingunmatched
    • Order Picking/Packingunmatched
    • PHP Scripting Language (PHP Hypertext Preprocessor)unmatched
    • Python Programming/Scripting Languageunmatched
    • Redisunmatched
    • Registrarunmatched
    • Reliability Engineeringunmatched
    • Replication and Remote Mirroringunmatched
    • Reporting Dashboardsunmatched
    • Right-Sizingunmatched
    • Root Cause Analysisunmatched
    • SSL-TLS (Secure Socket Layer - Transport Layer Security)unmatched
    • Single Sign-On (SSO)unmatched
    • Software Patchesunmatched
    • Sustainabilityunmatched
    • Systems Administration/Managementunmatched
    • Tableauunmatched
    • Team Playerunmatched
    • VPN (Virtual Private Network)unmatched
    • Vulnerability Scannersunmatched
    • Web Client Plug-insunmatched
    • Wikiunmatched
    • Willing to Travelunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder