Research Scientist - World Model

Luma
  • Redwood City, California
    30+ days ago

    Job Description

    THE ROLE
    This is the role at the center of the thesis. Luma already trains the strongest generative video models in the industry; the next step is turning those models into world models — interactive, controllable, physically faithful, and useful as a substrate for embodied reasoning. As a Research Scientist on the World Models team, you'll work on the next generation of generative models that can be rolled out as worlds.

    WHAT YOU'LL DO
    - Invent next-generation world model architectures — diffusion, transformer, autoregressive, or hybrid — with a particular focus on controllability and physical consistency.
    - Develop controllability mechanisms that let an agent step into the world: action conditioning, view conditioning, long-horizon rollouts.
    - Define and own the metrics: physical fidelity, long-horizon coherence, action-following, and downstream usefulness for policy training.
    - Run scaling studies that tell us where compute, data, and architecture pay off.
    - Publish at the frontier; contribute to the open-source release that is the long-term deliverable.

    MINIMUM QUALIFICATIONS
    - PhD or equivalent research record in ML, computer vision, robotics, or related fields.
    - Deep expertise in at least one of: large-scale generative modeling (video/3D/world), self-supervised representation learning, model-based RL.
    - Strong PyTorch and large-scale training experience — you've trained models that hit the limits of a multi-node cluster.
    - A research record the field knows (top-venue publications and/or widely-used open releases).

    PREFERRED
    - Prior work on world models, model-based RL, generative video, neural simulation, or 4D scene representations.
    - Experience using generative models for downstream embodied tasks (planning, control, evaluation).
    - Excitement about open-sourcing frontier models.

    About Luma


    Luma’s mission is to build unified general intelligence that can generate, understand, and operate in the physical world.

    We believe that multimodality is critical for intelligence. To go beyond language models and build more aware, capable and useful systems, the next step function change will come from vision. So, we are working on training and scaling up multimodal foundation models for systems that can see and understand, show and explain, and eventually interact with our world to effect change.

    Numbers & Facts

    LocationRedwood City, California

    Skills

    • Computer Visionunmatched
    • Metricsunmatched
    • Modeling Languagesunmatched
    • Open Sourceunmatched
    • Product/Service Launchunmatched
    • Publicationsunmatched
    • Roboticsunmatched
    • Scientific Researchunmatched
    • Simulationunmatched
    • Training/Teachingunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder