Experience with AI/ML infrastructure, multi-node GPU clusters, accelerated compute, model training or inference platforms, GPU scheduling, device plugins, Karpenter, cluster autoscaling, CUDA, NCCL, RoCE, InfiniBand, RDMA, SmartNIC/DPU offload, or high-performance AI/HPC networking is a significant plus. The OKE team owns a highly available 24x7 cloud service and is expanding the platform to support larger clusters, higher scale, improved operability, deeper OCI integrations, and increasingly demanding cloud native, AI, and GPU workloads.