D. in Computer Science, Computer Engineering, or equivalent practical experience Experience in any of the below is preferred: Proficiency with one or more modern ML frameworks (PyTorch, JAX, or TensorFlow), particularly the data loading and dataset access layer Columnar and lakehouse formats: Parquet, Iceberg, Delta, or Lance Distributed data loading frameworks for ML: Ray Data, NVIDIA DALI, WebDataset, or Mosaic StreamingDataset Performance engineering for I/O-bound workloads - Arrow, zero-copy, memory mapping, async I/O High-throughput object storage access patterns at GPU scale Data lineage and governance systems (DataHub, OpenLineage, Unity Catalog, or equivalent) Contributions to or operational experience with Spark, Daft, Polars, or DuckDB internals Containerization and orchestration technologies (Docker, Kubernetes). transformers, diffusion, retrieval-augmented generation) Proven experience building and delivering data and machine learning infrastructure in real-world production environments Familiarity with fine-tuning workflows, model optimization, and preparing models for scalable inference Familiarity with generative AI and its applications in accelerating and enhancing machine learning workflows Experience configuring, deploying and troubleshooting large scale production environments Experience in designing, building, and maintaining scalable, highly available systems that prioritize ease of use Extensive programming experience in Java, Python or Go Strong collaboration and communication (verbal and written) skills Comfortable navigating ambiguity and evolving technical landscapes, especially in fast-moving areas B.S., M.S., or Ph.