Senior Data Infrastructure Engineer

AirCall
  • San Francisco Office, CA
    4 days ago

    Job Description

    About the role

    Aircall's Data team is mid-migration: we are moving off a single Redshift cluster onto an Apache Iceberg lakehouse on S3, with Flink CDC into Kafka for ingestion and dbt-on-Spark via Apache Kyuubi on EKS for transformation. It's a real greenfield platform build - already scoped and underway - on top of a stack that carries ten years of startup-growth history, and all the quirks that come with it.

    We're building this role to give platform work the runway it deserves. Right now, our engineers wear two hats - owning the infrastructure and the business datasets running on top of it - and we're ready to invest in the high-leverage frameworks that will make both jobs easier: data quality automation, schema registry, and staging and gated promotion. This is a dedicated platform seat: your chance to build those foundations from the ground up. Your customers are the analytics engineers, data scientists, and AI agents who build on what you ship, and your product is their leverage.

    What you will do

    • Build and operate the lakehouse: Apache Iceberg on S3, table design and maintenance, partitioning and compaction, and the migration of remaining Redshift workloads onto it
    • Own ingestion end to end- Flink CDC Kafka (MSK) Iceberg, plus Rudderstack, Fivetran and DMS sources - and hold the freshness and reliability SLAs on it
    • Run and evolve the compute and orchestration layer: Apache Kyuubi on EKS for dbt-spark, Airflow (completing its ECS EKS migration), autoscaling, spot strategy and cost efficiency
    • Build the tooling, libraries and templates that let analytics engineers and data scientists own their own pipelines without filing a ticket - self-service is the deliverable, not a side effect
    • Close our environment gaps: a real staging environment, CI that tests against staging rather than production, automated schema-change detection, gated promotion and canary deploys for critical models
    • Own governance and access at the platform level: Lake Formation row/column RBAC, StrongDM zero-trust access, SSO, audit logging, and PII handling
    • Own observability: Monte Carlo, lineage, alerting and the SLAs we publish - and drive incidents to root cause and to a durable fix
    • Champion infrastructure as code and automation (Terraform, GitLab CI, GitOps) across everything the team runs

    What you own vs. our Analytics Engineers

    You own the platform: ingestion, storage, orchestration, compute, access control, observability and the frameworks on top of them. Our Analytics Engineers own the business-facing layer - dbt models, golden datasets, metric definitions and the semantic layer - and consume your platform as a service. You're energized by multiplying other people's speed, and drawn to problems where the end user is a fellow engineer.

    Must-haves

    • 4+ years (Senior: 6+) in data engineering, data platform or infrastructure engineering
    • Strong Python and SQL, with demonstrated experience building frameworks and tooling others depend on, not only pipelines
    • Production experience with an orchestration framework (Airflow, Dagster, Prefect) at meaningful scale - including the operational side, not just DAG authoring
    • Hands-on Apache Spark and distributed-systems fundamentals
    • Deep AWS experience (S3, EKS/ECS, IAM, Glue/Athena or equivalent)
    • Comfortable building and debugging CI/CD, infrastructure as code (Terraform) and GitOps workflows; familiar with Kubernetes and Docker
    • Track record owning reliability: SLAs, monitoring, alerting, on-call, and post-incident hardening
    • Daily, hands-on use of AI coding tools (Claude Code, Cursor, or equivalent) as a core part of how you build and operate infrastructure
    • Great cross-functional communication - you'll shape data contracts with backend engineering and align expectations with analytics consumers.

    Nice-to-haves

    • Production experience with an open table format (Apache Iceberg, Delta Lake, Hudi) and lakehouse migration off a classic warehouse
    • Streaming experience: Kafka/MSK, Flink, CDC pipelines, Kinesis
    • Experience with data governance tooling- Lake Formation, Unity Catalog, or equivalent RBAC/masking implementations
    • Familiarity with dbt (as a platform provider - dbt-spark, adapters, CI for dbt) and with data observability tooling such as Monte Carlo
    • Experience designing platforms consumed by AI/LLM workloads and low-latency analytics engines
    • Open-source contributions to data infrastructure projects

    Numbers & Facts

    LocationSan Francisco Office, CA

    Skills

    • Amazon Simple Storage Service (S3)unmatched
    • Amazon Web Services (AWS)unmatched
    • Apacheunmatched
    • Apache Kafkaunmatched
    • Apache Sparkunmatched
    • Artificial Intelligence (AI)unmatched
    • Artificial Intelligence (AI) Agentsunmatched
    • Automationunmatched
    • Autoscalingunmatched
    • Centers for Disease Control and Prevention (CDC)unmatched
    • Cisco Unityunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • Cross-Functionalunmatched
    • Data Scienceunmatched
    • Debugging Skillsunmatched
    • Distributed Computingunmatched
    • Dockerunmatched
    • Machine Toolunmatched
    • On Callunmatched
    • Open Sourceunmatched
    • Product Shipmentsunmatched
    • Python Programming/Scripting Languageunmatched
    • SQL (Structured Query Language)unmatched
    • Service Level Agreement (SLA)unmatched
    • Side Effectsunmatched
    • Single Sign-On (SSO)unmatched
    • Warehousingunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder