Build the annotation and auto-labeling pipelines that produce 3D-grounded supervision at scale, camera-relative 3D boxes, referring and spatial question answering, free space and reachability, ego-, world-, and object-centric reference frames, cross-view correspondence, camera motion, distance and size, and chain-of-thought traces, validated by programmatic and model-based critics. Build and operate the 3D and spatial evaluation suite, public benchmarks such as CV-Bench, BLINK, RefSpatial, VSI-Bench, SPAR-Bench, and RoboSpatial, NVIDIA's VANTAGE-Bench for real-world fixed-camera video understanding, and in-house benchmarks you design with continuous evaluation and full traceability from every reported score back to the exact weights, inputs, configuration, and evaluation code.