Lead Software Engineer, Fleet Management - DGX Cloud

NVIDIA Corp

Santa Clara, CA

JOB DETAILS
SKILLS
AngularJS, Application Programming Interface (API), Artificial Intelligence (AI), Automation, Autonomous Driving Systems, Business Processes, Cloud Computing, Data Management, Data Sets, Data Storage, Data Warehousing, Debugging Skills, Distributed Computing, Fleet Management, GPU (Graphics Processing Unit), JavaScript, JavaScript Frameworks, Leading Edge Technology, Linux Operating System, Network Operations Center, PostgreSQL, Problem Solving Skills, Programming Languages, Python Programming/Scripting Language, REST (Representational State Transfer), React.js, Sales Pipeline, Scalable System Development, Software Design, Software Engineering, Software as a Service (SaaS), System Architecture, System Integration (SI), Team Lead/Manager, Team Player, Technical Leadership, Technical/Engineering Design, Telemetry
LOCATION
Santa Clara, CA
POSTED
30+ days ago

Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent. As an NVIDIAN, you'll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. NVIDIA is widely recognized as one of the most desirable employers, with some of the most talented people in the world working for us. If you're passionate about building scalable, efficient systems to power cloud operations, we invite you to join our team.

We are looking for a Lead Software Engineer to join our DGX Cloud team and build the foundational systems that drive NVIDIA's high-performance GPU infrastructure. You will play a technical lead role in designing scalable cloud services that integrate with diverse systems including GPU telemetry in datacenters, and enabling operational automation across global cloud operations.

What You'll Be Doing:

  • Act as technical lead for a team of software engineers designing cloud services backed by databases and data warehouses.

  • Design and develop RESTful APIs to ingest telemetry from AI datacenters.

  • Build scalable cloud services for high-volume ingestion, processing, and storage of large datasets.

  • Build and manage data pipelines for online and offline data storage.

  • Collaborate across teams to codify business processes into scalable, self-measuring systems.

  • Optimize the reliability and efficiency of cloud services and operations.

  • Lead and ship impactful technical projects, ensuring quality and scalability at every stage.

What We Need To See:

  • At least 12+ years of industry experience with a Bachelor's or Master's degree (or equivalent experience); PhD degree preferred.

  • Expertise in building scalable REST APIs backed by PostgreSQL-compatible data stores. Proficiency in programming languages such as Go or Python.

  • Familiarity with modern JavaScript frameworks (e.g., React, Angular, Next.js).

  • Expertise in cloud infrastructure (AWS, GCP, Azure, etc) and container technologies like Docker and Kubernetes.

  • Expertise with high-scale distributed systems, including architectural patterns for APIs and data pipelines.

  • Outstanding communication and collaboration skills, with a focus on solving complex operational challenges.

  • A passion for delivering scalable and efficient cloud services.

  • Familiarity with Linux operating systems.

Ways to Stand Out from the Crowd:

  • A track record of leading engineers to successful delivery and operations of high-performance cloud services at Internet scale.

  • Experience operating NVIDIA datacenter GPUs.

  • Strong debugging and problem-solving skills in distributed environments.

NVIDIA is committed to creating an environment where diverse perspectives drive innovation. As part of the DGX Cloud team, you'll work on cutting-edge technology that powers the future of AI and cloud computing.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until April 24, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

About the Company

N

NVIDIA Corp

Visualize your future . . . We Do
NVIDIA is the world leader in graphics processing technologies, creating innovative, industry-changing products for computing, consumer electronics, and mobile devices. NVIDIA products are transforming visually-rich applications such as video games, film production, broadcasting, industrial design, space exploration, and medical imaging. We invest in our people and our technologies, support and fund industry research around the world, and consistently deliver high-quality products. NVIDIA's culture promotes and inspires a team of world-class employees to be at the top of their game. We've created an environment where talents are recognized and collaboration is valued. Our employees are shaping the world of tomorrow. . . today. We invite you to explore the opportunities available at NVIDIA to see what your future may hold.

COMPANY SIZE
10,000 employees or more
INDUSTRY
Computer Software
FOUNDED
1993
WEBSITE
http://www.nvidia.com