Able to write, optimize, and debug complex queries for large-scale analysis and transformation, including complex joins, window functions, and common table expressionsStrong hands-on experience with Apache Spark, processing large datasets efficiently across distributed clustersHands-on experience with the Databricks platform, including workspace management, notebook collaboration, and Delta Lake optimizationStrong proficiency with PythonDemonstrated experience building, monitoring, and maintaining robust end-to-end ETL and ELT pipelines in productionSolid understanding of data quality and data governance practice, including creating data quality rules and implementing anomaly detection to protect data integrity and consistencyWorking knowledge of data modeling and turning raw data into reusable analytics datasets and data productsExperience integrating or consuming data through REST APIsExperience with Git-based development workflows and modern software engineering practicesComfort working in terminal and command-line environmentsAbility to proactively engage subject matter experts, decode complex business logic, and turn domain knowledge into technical data requirementsStrong communication skills, including the ability to explain technical data concepts to non-technical stakeholders and deliver solutions that meet a stated business needDemonstrated ability to turn messy data into structured, high-value assets that support informed decision-makingAbility to manage multiple data workstreams at once, navigate ambiguity, and drive assigned work to resolution without close supervisionBachelor's degree in Computer Science, Engineering, Data Science, or a related technical fieldU.S. You will work directly in Spark and Databricks to land data from clinical and operational source systems, transform it into curated silver and gold datasets, and make those datasets reliable enough that mission users trust them without checking your work.