You bring experience:- Building and shipping production data pipelines in Python, with an understanding of what makes them reliable and maintainable under real load- Working with PySpark or a comparable distributed processing framework on datasets too large for a single machine- Writing SQL well enough to answer hard questions about data, not just retrieve it- Working in an operationally-driven environment where reliability and on-time delivery matter as much as new feature work- Working directly with non-engineering partners, subject matter experts, analysts, or customer-facing teams, and communicating clearly about data with people who don't read code- Holding a high bar in code review and expecting the same from those who review your work- Identifying data problems early and seeing work through to resolution rather than handing it off REQUIREMENTS - 3+ years of experience in software or data engineering, with meaningful Python in your background- Demonstrated experience building and maintaining production-grade data pipelines in Python- Hands-on experience with PySpark or a similar distributed data processing framework- Strong SQL skills, including working with large, messy, multi-source datasets- Strong understanding of software quality practices: testing, code review, documentation, and CI/CD- Experience working with cross-functional and non-technical stakeholders- Experience with pipeline orchestration tooling (Argo, Airflow, Databricks, dbt, or similar) preferred- Familiarity with clinical trial data, healthcare data, or another regulated data domain a plus- Familiarity with entity matching or data mastering a plus- Familiarity with AWS services (S3, Lambda, ECS, or similar) a plus COMPENSATION This role pays $110,000 to $135,000 per year, based on experience, in addition to stock options. Develop the transformation logic that maps raw trial and customer data to H1's internal data models, handling diverse source formats including CSV, JSON, Parquet, and APIs.- Write and tune SQL against large datasets to investigate data questions, validate pipeline output, and support analysis that clinical SMEs and customer-facing teams depend on.- Turn around customer-driven changes quickly, scoping requests as they arrive, shipping changes that hold up under enterprise SLAs, and reworking logic as customer needs shift mid-flight.-