You'll engage hands-on across the entire customer lifecycle: leading new GPU cluster bring-up and acceptance, driving InfiniBand/RoCE fabric validation and HPC performance benchmarking, defining how we operate customer bare-metal fleets at rack-level-and-up (IT service, break-fix, network, and firmware), and standing up locked-down, security-sensitive environments for our most strategic AI customers. Working alongside the teams that build and operate each layer, you are the deep technical expert who turns raw data center hardware—racks, GPUs, high-speed fabric, firmware—into reliable compute that customers can train and inference on at scale, spanning infrastructure engineering, provisioning, validation, operations, and support.