Define and build the evaluation methodology for predictive world models, spanning representation quality, prediction accuracy, generation fidelity, physical plausibility, long-horizon rollout consistency, and ultimately closed-loop policy performance, then feed the resulting models back into training as a source of synthetic data and corner-case simulation. In this role, you will research, implement, and evaluate world models that learn the dynamics of the physical world from large-scale multimodal data — predicting how a scene evolves under an agent's actions, and serving as a learned simulator for training and evaluating driving and robotic policies.