Founding Data Engineer — AI Healthcare
The Opportunity
Our client is building an AI-powered healthcare platform designed to make high-quality care more accessible at massive scale.
They’re now hiring their first dedicated Data Engineer to build the data foundation behind the company.
This is not a role where you inherit a mature platform and optimize around the edges.
You’ll own how data moves through the business from end to end — from production systems into the lakehouse and warehouse, through transformation and governance, and ultimately into the hands of AI, product, finance, partnerships, and leadership.
If you’ve wanted the opportunity to define how a company thinks about data from the ground up, this is it.
What You'll Own
Build the Data Platform
Design and operate reliable CDC and ELT pipelines from MariaDB, PostgreSQL, and MongoDB into S3, Apache Iceberg, and Snowflake
Create a governed, trusted source of production data that the entire company can build on
Design a scalable warehouse architecture with clean raw, transformed, and business-ready layers
Implement monitoring, alerting, and reliability standards across the data stack
Create Trusted Business Data
Build the transformation layer using dbt or similar tooling
Turn raw production data into tested, documented, version-controlled models
Establish trusted definitions for metrics such as:
Visits
Bookings
Revenue
Retention
Product engagement
Power executive reporting and downstream analytics from a consistent source of truth
Own Orchestration & Reliability
Select and implement the right orchestration platform for the company
Automate pipelines, transformations, and dashboard refreshes
Build monitoring and alerting so failures are caught quickly
Establish reliability standards as the volume and complexity of the platform grows
Build Healthcare-Grade Data Governance
You’ll play a critical role in determining how sensitive healthcare data is handled.
That includes:
Row- and column-level access controls
PHI restrictions
HIPAA-aligned data architecture
Safe Harbor anonymization
Data deletion workflows
Role-based access policies
Secure datasets for analytics and AI use cases
The goal is to make data highly useful without compromising patient privacy or security.
Enable AI & Product Teams
Build datasets and pipelines supporting AI model training and evaluation
Partner with AI engineers on training data and data quality
Support product teams with trustworthy behavioral and product data
Help finance, marketing, partnerships, and leadership answer important business questions without creating separate versions of the truth
What We're Looking For
5+ years of data engineering experience
Experience owning production data infrastructure end to end
Strong SQL and Python
Experience building and maintaining CDC / ELT pipelines
Familiarity with tools such as Fivetran, Airbyte, or similar platforms
Hands-on experience with modern warehouse or lakehouse architectures
Experience with:
AWS S3
Apache Iceberg
Snowflake or similar data warehouses
Data catalogs
dbt or comparable transformation frameworks
Experience with orchestration platforms such as Airflow, Dagster, or AWS Glue
Strong AWS fundamentals including IAM, Lambda, Kinesis, and Glue
Strong understanding of production reliability and data quality
The Type of Engineer Who Thrives Here
This role is best suited for someone who:
Likes building systems from scratch
Doesn't need a perfectly defined roadmap before getting started
Can evaluate tools rather than simply use whatever is already installed
Thinks about reliability, governance, and maintainability from day one
Can translate business questions into durable data models
Communicates well with technical and non-technical stakeholders
Wants meaningful ownership instead of narrowly scoped tickets
Enjoys being the person people turn to when the answer starts with, "What does the data actually say?"
Particularly Relevant Experience
Experience in any of the following would be especially valuable:
HIPAA, PHI, or healthcare data
Healthcare technology
Data anonymization and governance
ML training datasets and feature pipelines
SageMaker, Databricks, or Jupyter environments
ClickHouse or high-volume event pipelines
Server-side tracking, CDPs, or behavioral analytics
BI tooling such as Metabase
Semantic or metrics layers
First data engineer or early-stage startup experience
Healthcare experience is helpful, but the bigger requirement is that you've built reliable, governed data infrastructure in production.
Why This Role
You'll have an unusually broad mandate.
Your work will directly influence:
How the company measures performance
How executives make decisions
How AI models are trained
How patient data is protected
How product teams understand behavior
How the company scales its analytics infrastructure
Instead of joining a large data organization and owning one piece of the stack, you'll have the opportunity to design the stack itself.
Compensation
Base Salary: $200,000 – $275,000
Equity: Meaningful ownership based on experience and level
This is an opportunity to become the technical owner of the data foundation behind a rapidly scaling AI healthcare company.
| Location | New York, NY |
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.
Free resume templatesImprove your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.
Free resume builder