Artificial Intelligence (AI), Audio System Design, Benchmarking, Communication Skills, Cross-Functional, Data Sets, Develop Methodologies, Machine Learning, Mathematics, Performance Metrics, Production Systems, Prototyping, Scientific Research, Startup, Team Player, Training/Teaching, Voice Applications
This is a research-driven, high-impact role for ML researchers who want to push the boundaries of real-time AI. As a Founding Machine Learning Research Engineer at Retell, you'll focus on advancing model capabilities for human-like voice agents operating in complex, real-world environments.
You'll explore new approaches across LLMs and audio models, design novel evaluation methods, and prototype systems that improve reasoning, latency, and conversational quality. Your work will directly influence production systems, bridging cutting-edge research with real-world deployment.
If you're excited about solving open-ended ML problems, experimenting rapidly, and shaping how voice AI systems think and perform, this is a unique opportunity to do so at scale.
KEY RESPONSIBILITIES
- Research & Experimentation – Explore and develop new techniques across LLMs and audio models to improve reasoning, latency, and conversational quality in real-time systems.
- Model Prototyping – Rapidly build and iterate on experimental models and pipelines, turning research ideas into working prototypes.
- Evaluation & Benchmarking – Design novel evaluation frameworks, datasets, and metrics to measure performance on complex, real-world voice tasks.
- Bridge Research to Production – Collaborate closely with engineering to translate research insights into deployable systems.
- Human Feedback Loops – Develop methods to incorporate human evaluation into model improvement, especially for subjective conversational quality.
- Advance the Frontier – Stay at the cutting edge of ML research and bring new ideas into Retell's product and infrastructure.
HOW TO THRIVE
- Strong ML Research Background – You've worked on advanced ML problems (for example: LLM pre-training and post training, transcription model training, text to speech model training, or multimodal systems), either in industry or academia.
- Deep Technical Foundation – Comfortable with PyTorch, model architectures, and the math behind modern machine learning.
- Experimental Mindset – You enjoy exploring open-ended problems and iterating quickly on ideas.
- Bridging Theory & Practice – You can translate research into systems that work in real-world environments.
- Startup-Ready – You thrive in fast-paced environments with high ownership and ambiguity.
- Collaborative & Clear Communicator – You can explain complex ideas and work cross-functionally to drive impact.
Tech stack:
PyTorch, LLMs, Audio/Speech Models, Text-to-Speech (TTS), Automatic Speech Recognition (ASR), Multimodal Systems, Python
Seniority:
1 - 5 years of experience in audio/multimodal AI research or engineering
Work experience:
Working at a high bar company with AI products (MAANG, vc backed startup, etc.)
Experience with LLM pre-training or post-training, evals, and translating research to production
Coming from another top voice ai or audio startup (Cartesia, Eleven Labs, Descript, etc.)
Education:
Degree in CS, ML, or closely related field (PhD prefered)
Recent publications in voice, audio, or speech AI
Hard skills:
Hands-on PyTorch and audio/speech model development
Experience with TTS, ASR, or multimodal audio systems
Pre-training experience at scale
Miscellaneous:
Comfortable with intense startup pace including weekend work