The Seed Multimodal Interaction and World Model team is dedicated to developing models that boast human-level multimodal understanding and interaction capabilities. The team also aspires to advance the exploration and development of multimodal assistant products.
Responsibilities:
Design and implement reinforcement learning (RL) training systems for large-scale multimodal foundation models
Develop unified modeling frameworks that integrate video, audio, and language, with a focus on visual latent reasoning
Explore RL-based approaches to bridge understanding and generation for multimodal visual reasoning
Collaborate with researchers to evaluate models on tasks involving world modeling, reasoning, and instruction-conditioned generationMinimum Qualifications:
Currently pursuing a PhD in Software Development, Computer Science, Computer Engineering, or a related technical discipline
Publications in accredited venues, such as CVPR, ECCV, ICCV, NeurIPS, ICLR, ICML, or other leading conferences in AI and ML
Strong research background in at least one of the following: reinforcement learning, multimodal learning, video understanding, or vision-language modeling
Preferred Qualifications:
Experience with reinforcement learning in multimodal or interactive environments
Familiarity with video generation or diffusion-based generative models
Experience with large-scale model training (e.g., distributed training, curriculum learning, or memory-augmented transformers)
Solid programming and engineering skills, with experience building training or evaluation pipelines for ML models
As a condition of employment, all successful candidates must be able to establish authorization to work in the United States. For this position, the Company does not provide sponsorship or any immigration-related benefits.
Numbers & Facts
Location
San Jose, CA
More jobs like this
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.