Build and operate the infrastructure that keeps large-model training and online inference running reliably.
You will support GPU clusters at 100+ card scale, respond to production incidents, and work closely with training and inference teams on stability, performance, and resource efficiency.
This is a hands-on infrastructure role across Linux, Slurm, Kubernetes, Ceph, RDMA networking, GPU drivers, observability, and automation.
Responsibilities
Provide infrastructure support for large-model training and online inference workloads, responding quickly to production incidents and operational issues.
Deploy and operate 100+ GPU clusters, ensuring training and inference jobs run stably at scale.
Maintain Slurm scheduling and Kubernetes platforms, optimizing resource allocation and multi-tenant isolation.
Deploy, expand, and tune Ceph distributed storage for large-scale AI workloads.
Operate RDMA networks such as InfiniBand and RoCE, plus GPU fleet components including DCGM, drivers, CUDA, and NCCL version management.
Build automation tools, monitoring, and alerting systems that improve cluster stability and reduce debugging time.
Requirements
Senior Linux operations background, with hands-on experience in GPU clusters or HPC environments.
Experience operating 100+ GPU clusters that support large-model training or online inference workloads.
Deep understanding of the Linux kernel, networking stack, and storage stack, with the ability to diagnose low-level performance bottlenecks.
Expertise with Slurm for training workloads and production Kubernetes clusters for inference, including architecture design and performance tuning.
Strong experience with Ceph distributed storage, including large-scale deployment, expansion, and performance tuning.
Familiarity with RDMA networking such as InfiniBand or RoCE, and the ability to help debug NCCL collective communication issues.
Strong Python and Shell scripting skills, with experience building automation platforms or operations tooling.
Experience with Prometheus and Grafana monitoring systems, including large-scale cluster monitoring and alerting.
Strong ownership for production reliability, including leading incident response, postmortems, and on-call responsibilities.
Nice to Have
Ability to partner with training teams to diagnose performance bottlenecks across NCCL communication and storage I/O.
Experience building GPU clusters from zero to one at 100+ to 1,000+ GPU scale.
Familiarity with additional distributed storage systems such as MinIO, Weka, or Lustre.
HPC operations experience and familiarity with parallel file systems.
Open-source contributions or technical writing related to infrastructure, HPC, GPU operations, or reliability.
How to Apply
Send your resume along with GitHub, personal project, or technical writing links to the contact below.
For open-source contributions or past projects, direct links or short write-ups are welcome.
Take-home tasks, if any, will be paid at a reasonable market rate.
No requirements around years of experience or degree - we evaluate on technical depth and past work.
Numbers & Facts
Location
San Francisco, CA
Skills
Artificial Intelligence (AI)unmatched
Automationunmatched
CUDA (Compute Unified Device Architecture)unmatched
Debugging Skillsunmatched
Device Driversunmatched
Distributed Computingunmatched
Energy Efficiencyunmatched
File Systemsunmatched
GPU (Graphics Processing Unit)unmatched
GitHubunmatched
Home Automationunmatched
Identify Issuesunmatched
Incident Responseunmatched
Input/Outputunmatched
Large-Scale Systemsunmatched
Linux Kernelunmatched
Linux Operating Systemunmatched
Machine Toolunmatched
On Callunmatched
Online Trainingunmatched
Open Sourceunmatched
Performance Tuning/Optimizationunmatched
Python Programming/Scripting Languageunmatched
Resource Managementunmatched
Return on Capital Employed (ROCE)unmatched
Technical Analysisunmatched
Technical Writingunmatched
Time Managementunmatched
Truck Driverunmatched
Unix Shell Programmingunmatched
Vehicle Fleetsunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.