Own the offline dataset pipeline - design, implement, test, and deploy Cloud-based pipelines that convert logged multi-sensor data into VLM/VLA training datasets, spanning geometric labels (3D/2D detection, tracking, segmentation, depth) through semantic, scenario-level, and action/trajectory-grounded annotations. Sitting within Offline Perception, this team turns petabytes of logged multi-modal fleet data (images, kinematics) into VLM/VLA-ready datasets: geometric annotations, scenario-level semantic descriptions, action- and trajectory-grounded labels, and reasoning traces that explain why a maneuver was taken.