You''ll engage hands-on across the entire customer lifecycle: leading new GPU cluster bring-up and acceptance, driving InfiniBand/RoCE fabric validation and HPC performance benchmarking, defining how we operate customer bare-metal fleets at rack-level-and-up (IT service, break-fix, network, and firmware), and standing up locked-down, security-sensitive environments for our most strategic AI customers. Working alongside the teams that build and operate each layer, you are the deep technical expert who turns raw data center hardware-racks, GPUs, high-speed fabric, firmware-into reliable compute that customers can train and inference on at scale, spanning infrastructure engineering, provisioning, validation, operations, and support.