Hands-on experience designing, deploying, operating, or troubleshooting high-performance GPU networks in on-premises or cloud environments using technologies such as InfiniBand, RoCE, or GPUDirect RDMA. Profiling distributed training and inference workloads to identify bottlenecks across GPUs, CPUs, memory, network fabrics, storage systems, and software stack.