Backend Engineer — Data Platform

Savant Bio
  • New York
    7 days ago

    Job Description

    About Us

    Savant is transforming how healthcare and life sciences organizations unlock the value trapped in unstructured medical data. Our platform combines cutting-edge large language models (LLMs) with domain-specific quality controls to convert free-text clinical records into structured, analysis-ready data — efficiently, accurately, and at scale. We work with leading institutions across healthcare, life sciences, and research, supporting faster studies, sharper insights, and better care. Backed by Roivant (NASDAQ: ROIV), Savant is built for organizations that see structured data not just as an output, but as a foundation for innovation.

    The Opportunity

    We’re hiring a Backend Engineer to help mature and scale the data platform underlying Savant’s clinical data products. You’ll help design and implement the systems that move large, complex datasets through ingestion, processing, quality control, and delivery.

    This is a backend engineering role with a strong data engineering and infrastructure orientation.

    You’ll work across Python and Rust services, SQL-based data systems (DuckDB, Postgresql, and Ducklake), Dagster pipelines, and Kubernetes workloads. You’ll help establish the architectural patterns that allow Savant to process sensitive healthcare data efficiently while maintaining reproducibility, traceability, and quality.

    Savant is based in NYC, and this role is remote-eligible. Preference will be given to candidates able to come to work in NYC.

    Some Things You Might Work On

    • Design and build production pipelines for ingesting, transforming, validating, and delivering healthcare datasets
    • Develop backend services and APIs that support Savant’s clinical data products
    • Build reliable Dagster workflows with strong patterns for partitioning, retries, idempotency, backfills, and observability
    • Design data models and contracts that support schema evolution, lineage, reproducibility, and downstream analysis
    • Operate and improve containerized workloads running on Kubernetes
    • Build systems for monitoring pipeline health, data quality, performance, and infrastructure cost
    • Diagnose production issues across application, data, and infrastructure layers
    • Improve the performance of Python and SQL workloads operating over large datasets
    • Establish testing and deployment practices that make data systems safer and easier to change
    • Collaborate with AI engineers to productionize new model workflows and scale them across customer datasets
    • Help shape technical architecture as Savant’s platform, customer base, and processing volume grow

    What We’re Looking For

    • 5+ years of experience building production backend or data-intensive software systems
    • Excellent Python and SQL skills
    • Experience with AWS and Kubernetes
    • Sound technical judgment, taste, and the ability to make pragmatic architectural decisions
    • Experience designing and operating data pipelines at scale using Dagster, Airflow, Prefect, or a comparable orchestration framework
    • Clear communication skills and a collaborative mindset

    Bonus Points For

    • Experience processing clinical, claims, laboratory, registry, or other healthcare data
    • Experience with Dagster
    • Experience with distributed batch processing or compute-intensive data pipelines
    • Familiarity with HIPAA, de-identification, auditability, or other information security requirements for sensitive data
    • Experience supporting machine-learning or LLM inference pipelines
    • Experience in an early-stage startup or small, fast-moving technical team

    Why Join Us

    • Foundational Ownership: Build the systems on which Savant’s products and AI workflows depend
    • Meaningful Scale: Solve real data-engineering and operational challenges involving large, complex clinical datasets
    • Broad Technical Scope: Work across backend services, data architecture, orchestration, and infrastructure
    • Tight Collaboration: Work directly with AI, product, clinical, and engineering partners
    • Mission-Driven Work: Help make clinical data more usable for research, decision-making, and patient care

    Base salary for this role will be determined during the interview process and will vary based on multiple factors, including but not limited to prior experience, relevant expertise, current business needs, and market conditions. The expected base salary for the role will generally be between $175,000 to $250,000 per year, with meaningful equity. The final salary offered may be outside this range based on individual circumstances and business and market conditions.

    Base salary if hired is only part of the total compensation package, which, depending on the position, may also include other components such as discretionary bonuses and Company-sponsored benefit programs.

    This position is at-will and Savant reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance and business and market conditions.

    Savant provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

     

    Numbers & Facts

    LocationNew York

    Skills

    • Amazon Web Services (AWS)unmatched
    • Analysis Skillsunmatched
    • Application Programming Interface (API)unmatched
    • Architectural Servicesunmatched
    • Artificial Intelligence (AI)unmatched
    • Biologyunmatched
    • Claims Processingunmatched
    • Clinical Dataunmatched
    • Communication Skillsunmatched
    • Compensation and Benefitsunmatched
    • Data Managementunmatched
    • Data Modelingunmatched
    • Data Processingunmatched
    • Data Qualityunmatched
    • Data Setsunmatched
    • HIPAA (Health Insurance Portability and Accountability Act)unmatched
    • Healthcareunmatched
    • Identify Issuesunmatched
    • Information/Data Security (InfoSec)unmatched
    • Machine Learningunmatched
    • Medical Productsunmatched
    • Medical Recordsunmatched
    • Modeling Languagesunmatched
    • Patient Careunmatched
    • Performance Managementunmatched
    • PostgreSQLunmatched
    • Python Programming/Scripting Languageunmatched
    • Quality Assurance Methodologyunmatched
    • Quality Controlunmatched
    • SQL (Structured Query Language)unmatched
    • Scientific Researchunmatched
    • Startupunmatched
    • Structured Analysisunmatched
    • Structured Dataunmatched
    • Team Playerunmatched
    • Technical Recruitingunmatched
    • Traceabilityunmatched
    • Training Data Setsunmatched
    • Unstructured Dataunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder