AI Workload Expertise: Hands-on experience running real AI compute jobs at scale - pre-training, fine-tuning, or large-scale inference of open-source or proprietary models - including practical familiarity with distributed training strategies (data, tensor, pipeline, and expert parallelism) and frameworks such as PyTorch, Megatron-LM, DeepSpeed, or equivalent. Lead deep diagnosis of large-scale cluster failures and performance regressions, isolating root cause across the full stack: GPU and NIC firmware, PCIe/NVLink topology and NUMA placement, InfiniBand/RoCE fabric health, congestion control and routing, storage and data-loader throughput, scheduler placement, and framework/communication-library behavior.