Contribute to pre-training data curation for LLM development across all domains and modalities (eg: web, STEM, code, multilingual, image, audio, video) Design, implement, and deploy scalable ML models for information extraction, data selection, and synthetic data generation over trillions of unstructured records Design, execute, and analyze scientific experiments to advance our understanding of large language models Independently lead small-scale research projects while contributing to cross-functional, larger-scale research initiatives Accelerate workflows by building high-quality data tooling that improves data velocity, reliability, and scalability Apply expertise in one or more of the following areas: multimodal generation and perception (text, image, video, or audio), OCR, data scaling laws, or data mixing. 4+ years of experience with large language models (LLMs), large multimodal models (LMMs), computer vision, or related AI/ML technologies Experience developing and evaluating machine learning models, with a strong understanding of data and model quality Strong programming skills and hands-on experience using one or more deep learning frameworks, such as PyTorch, TensorFlow, or JAX Experience building large-scale machine learning systems and working with distributed data processing frameworks Strong problem-solving skills with a results-oriented mindset Experience leading rapid prototyping and proof-of-concept development for AI/ML applications Excellent communication skills, with the ability to work independently and collaborate effectively with cross-functional teams Masters degree or Ph.