Hiring Research Scientists to help define and advance the core technical direction of an early-stage AI company focused on multimodal representation learning.
This role is for a researcher who has trained models natively across multiple modalities, audio, images, video, text, and other structured inputs, and understands how richer representations can improve downstream reasoning and action. You’ll work directly with the founder, help set research direction from the earliest stage, and have significant autonomy over what gets explored and built.
What You’ll Own
Lead research in multimodal representation learning
Train and improve models across audio, image, video, text, and related modalities
Explore architectures that learn unified representations across multiple input types
Help define research priorities and long-term technical direction
Design experiments, evaluate model behavior, and identify promising research paths
Work across ambiguous, greenfield research problems with substantial autonomy
Translate research insights into systems that can ultimately support real-world AI products
What We’re Looking For
3–6 years of relevant research experience, with flexibility for exceptional senior candidates
Deep expertise in multimodal representation learning
Track record of training models natively across multiple modalities
Strong understanding of modern deep learning architectures, training techniques, and evaluation
Experience conducting frontier-level research at a leading AI lab, research organization, or similarly high-bar environment
Research judgment strong enough to independently propose and drive new technical directions
Comfortable working in a very early-stage, high-intensity environment with significant ambiguity
High ownership and the ability to operate without a predefined research roadmap
Strong Green Flags
Direct experience training omni-models or native multimodal models
Research spanning combinations of video, audio, vision, language, and structured signals
Experience at leading multimodal or video-generation organizations such as Luma, Runway, Pika, or comparable frontier AI teams
Published research in multimodal learning, representation learning, generative modeling, video models, or adjacent fields
Prior research or technical leadership experience
Experience taking research ideas from hypothesis through large-scale training and rigorous evaluation
Exceptional technical or competitive achievement, including IOI, IMO, quantitative research, or other world-class competitive backgrounds
The strongest candidate has worked directly on models that learn across several modalities rather than simply combining independently trained models at inference time. You should understand the challenges of cross-modal representation learning, model training, conditioning, data quality, evaluation, and scaling—and be interested in pushing those systems further.
Strong multimodal researchers from adjacent areas, particularly video generation and multimodal foundation models, are also highly relevant.
Numbers & Facts
Location
San Francisco, California
Skills
Artificial Intelligence (AI)unmatched
Data Analysisunmatched
Data Qualityunmatched
Deep Learningunmatched
Experiment Designunmatched
Leadershipunmatched
Quantitative Researchunmatched
Research Laboratoryunmatched
Research Skillsunmatched
Scientific Researchunmatched
Technical Leadershipunmatched
Technical Researchunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.