You will guide work across Kubernetes clusters, accelerator enablement, workload scheduling, high-performance networking, storage, and capacity management, partnering with AI researchers and engineers to improve reliability, performance, scalability, and developer productivity. Familiarity with frameworks such as PyTorch, Ray, vLLM, SGLang, or TensorRT-LLM, or experience qualifying accelerators and tuning distributed workloads.