GPU Network Engineer
Position Overview
I'm partnering with a rapidly growing AI infrastructure company building and operating large-scale GPU environments that support AI training, fine-tuning, and inference workloads.
This is an opportunity to design, build, and scale the high-performance network fabrics that power next-generation AI infrastructure. You'll focus on the networking layer that enables GPU clusters to communicate at scale, working across InfiniBand, RoCEv2, and high-speed Ethernet environments to deliver low-latency, lossless connectivity for mission-critical AI workloads. Working closely with infrastructure, compute, and operations teams, you'll play a key role in supporting the performance, scalability, and reliability of advanced AI platforms.
Key Responsibilities
- Design, deploy, and support high-performance east-west networking for GPU-to-GPU, rack-to-rack, and cluster-to-cluster communication.
- Build and operate InfiniBand, RoCEv2, and high-speed Ethernet fabrics supporting distributed AI workloads.
- Implement and maintain network architectures including CLOS/ECMP, EVPN/VXLAN, MLAG, and other high-availability designs.
- Support routing, traffic engineering, resiliency, and connectivity across data center and multi-site environments.
- Automate network provisioning, monitoring, and troubleshooting using modern infrastructure and observability tools.
Qualifications
Required
- 5+ years of experience supporting data center, HPC, or AI infrastructure networking environments.
- Hands-on experience designing and operating GPU cluster networks supporting GPU-to-GPU, rack-to-rack, and cluster-to-cluster communication.
- Deep experience with InfiniBand (NDR/XDR) and RoCEv2 networking technologies.
- Experience with NVIDIA networking technologies, including UFM, SHARP, and HPC fabric best practices.
- Knowledge of CLOS/ECMP architectures and east-west traffic optimization.
- Experience with EVPN/VXLAN and high-speed Ethernet environments.
- Experience supporting networking for bare-metal, virtualized, and Kubernetes-based environments.
- Experience with network automation and infrastructure tools such as NetBox, Netconf, Ansible, or Terraform.
- Hands-on experience with Arista EOS, Juniper Junos, NVIDIA/Mellanox, or similar networking platforms.
- Strong troubleshooting, packet analysis, and network performance tuning skills.
Nice to Have
- Experience tuning InfiniBand or RoCEv2 fabrics for large-scale AI training environments.
- Familiarity with congestion management, adaptive routing, and lossless fabric design.
- Understanding of AI workload characteristics across training, fine-tuning, and inference environments.
- Experience with Python-based network automation or software-defined networking technologies.
- Networking certifications such as CCNP, CCIE, JNCIP, JNCIE, or equivalent.
Benefits
- $185,000 to $225,000 base salary
- Annual performance bonus
- Restricted Stock Units (RSUs)
This opportunity is ideal for a network engineer who enjoys building high-performance infrastructure at scale. You'll have the opportunity to solve complex networking challenges, work with cutting-edge AI compute environments, and play a critical role in enabling the next generation of AI applications through world-class GPU networking infrastructure.
- For this position, you must be currently authorized to work in the United States without the need for sponsorship for a non-immigrant visa. CyberCoders will consider for Employment in the City of Los Angeles qualified Applicants with Criminal Histories in a manner consistent with the requirements of the Los Angeles Fair Chance Initiative for Hiring (Ban the Box) Ordinance.This job was first posted by CyberCoders on 08/19/2026 and applications will be accepted on an ongoing basis until the position is filled or closed.
Everforth CyberCoders is proud to be an Equal Opportunity Employer All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, age, sexual orientation, gender identity or expression, national origin, ancestry, citizenship, genetic information, registered domestic partner status, marital status, status as a crime victim, disability, protected veteran status, or any other characteristic protected by law. Our hiring process includes AI screening for keywords and minimum qualifications, and a virtual recruiter as part of the application process. A human recruiter reviews all results. Click here for details on our virtual recruiter . Everforth CyberCoders will consider qualified applicants with criminal histories in a manner consistent with the requirements of applicable state and local law, including but not limited to the Los Angeles County Fair Chance Ordinance, the San Francisco Fair Chance Ordinance, and the California Fair Chance Act. Everforth CyberCoders is committed to working with and providing reasonable accommodation to individuals with physical and mental disabilities. Individuals needing special assistance or an accommodation while seeking employment can contact a member of our Human Resources team at Benefits@CyberCoders.com to make arrangements.
Salary
$185000 - $225000 USD per year