Staff Software Engineer - Fleet Management

Nscale AS
  • NY
  • $220,000–$320,000 Per Year
30+ days ago

Job Description

.

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you''ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you''ll be contributing to building the technology that powers the future.

About the role

Nscale is hiring a Staff Software Engineer to build Fleet Manager - the workflow automation platform that provisions, tests, and remediates GPU nodes and network switches at scale.

This role sits at the intersection of distributed systems, infrastructure automation, and physical hardware. You''ll own domain-level architecture within Fleet Manager: Python-based systems that manage the entire operational lifecycle of our compute infrastructure, from initial device enrollment through multi-day burn-in testing to ongoing health monitoring and automated remediation. The problems are challenging and the stakes are high - the software you design and build determines how quickly and how reliably Nscale scales its GPU fleet to meet demand.

This is an opportunity to shape a foundational platform early, setting the patterns and standards that engineers across Fleet Manager build on.

What you''ll work on

  • Device provisioning and enrollment: automation that takes bare-metal GPU nodes and network switches from first power-on to production-ready - BMC configuration, DHCP reservations, and provisioning state machines.
  • Burn-in and validation: multi-day testing workflows that qualify hardware before it enters, and re-enters, the fleet.
  • Workflow orchestration: durable, event-driven state machines that span multiple days, survive crashes, resume from checkpoints, support human-in-the-loop approval gates, and let thousands of concurrent idempotent workflows run without stepping on each other.
  • GPU health monitoring and self-healing: detection, diagnosis, and automated remediation workflows that keep nodes healthy in production.
  • Network configuration: switch lifecycle automation and network state management across the fleet.
  • Integrations: keeping Fleet Manager consistent with datacenter inventory tooling (DCIM, NetBox), bare-metal provisioning systems (MAAS, Ironic, IPMI), credential stores, and monitoring infrastructure.
  • Observability: structured logging, metrics, distributed tracing, and tooling that lets operators troubleshoot effectively.

Responsibilities

  • Domain-level technical direction. Own the architecture for a major Fleet Manager domain - such as provisioning, validation, or remediation - influencing engineers across the team and adjacent squads.
  • Design and build production-grade automation. Implement device provisioning, burn-in testing, network configuration, and hardware health validation workflows in Python.
  • Engineer for reliability and auditability. Treat idempotency, resumability, checkpointing, retries, replay, and failure handling as first-class design concerns.
  • Integrate broadly. Connect Fleet Manager with datacenter infrastructure management systems, cloud orchestration platforms, and bare-metal provisioning tools.
  • Create leverage through standards. Establish shared patterns, libraries, conventions, and operational runbooks that other engineers build on.
  • Own production outcomes. Operate what you build with strong observability, alerting, incident response, and day-2 operational discipline.
  • Mentor and influence. Raise the bar through design reviews, implementation guidance, and operational best practices.
  • Use AI to accelerate delivery while maintaining architectural coherence.

Requirements

  • Extensive experience designing, building, and operating distributed systems in production, ideally in infrastructure automation, workflow tooling, or platform engineering.
  • Strong proficiency in Python - Fleet Manager is built entirely in Python.
  • Strong understanding of event-driven and workflow architecture, including reliable delivery, idempotency, retries, replay, and failure handling.
  • Track record of delivering automation systems from ambiguous requirements to production, with hands-on day-2 operations experience (monitoring, incident response, performance optimization).
  • Proven ability to lead ambiguous technical work across team boundaries and drive domain-level delivery through influence rather than formal authority.
  • You are driven by building distributed systems at scale, infrastructure reliability, scalability, security, and continuous improvement.
  • You use AI tools like Claude or Cursor as a core part of your development workflow to create leverage, increase quality, and accelerate delivery.
  • Excellent communication skills to build consensus with stakeholders, both internally and externally, in a fast-paced, high-agency environment.

Preferred

  • Experience with workflow orchestration tools like Temporal, Airflow, Prefect, or similar
  • Hands-on experience with infrastructure tooling: DCIMs, NetBox, OpenStack, or ERP systems
  • Bare-metal provisioning and automation: MAAS, Ironic, IPMI, PXE boot, or network automation
  • Experience building hardware lifecycle automation: provisioning, validation, testing, or remediation workflows
  • GPU infrastructure experience: health monitoring, burn-in testing, or cluster management
  • HPC and networking: datacenter topology, high-performance interconnects (InfiniBand, RoCE)
  • Deep knowledge of Kubernetes, Infrastructure as Code (Terraform, Pulumi), AWS, and GCP
  • Open-source contributions in infrastructure automation or cloud-native tooling

The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.

Salary Range

$220,000-$320,000 USD

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Numbers & Facts

LocationNY
Salary$220,000–$320,000 Per Year

Skills

  • Amazon Web Services (AWS)unmatched
  • Architectural Servicesunmatched
  • Artificial Intelligence (AI)unmatched
  • Automationunmatched
  • Automation Systemsunmatched
  • Best Practicesunmatched
  • Business Operationsunmatched
  • Business Supportunmatched
  • Cloud Computingunmatched
  • Communication Skillsunmatched
  • Consensus Building Skillsunmatched
  • Continuous Improvementunmatched
  • Cost Controlunmatched
  • DHCP (Dynamic Host Configuration Protocol)unmatched
  • Distributed Computingunmatched
  • ERP (Enterprise Resource Planning)unmatched
  • Fleet Managementunmatched
  • GCP (Good Clinical Practices)unmatched
  • GPU (Graphics Processing Unit)unmatched
  • Home Automationunmatched
  • IPMI (Intelligent Platform Management Interface)unmatched
  • Identify Issuesunmatched
  • Incident Responseunmatched
  • Leadershipunmatched
  • Machine Toolunmatched
  • Mentoringunmatched
  • Metricsunmatched
  • Network Administration/Managementunmatched
  • Network Configuration Managementunmatched
  • Network Operations Centerunmatched
  • Network Switchingunmatched
  • Network System Hardwareunmatched
  • Network Testingunmatched
  • Network Topologyunmatched
  • Open Sourceunmatched
  • Operations Managementunmatched
  • Performance Tuning/Optimizationunmatched
  • Production Systemsunmatched
  • Python Programming/Scripting Languageunmatched
  • Reliability Engineeringunmatched
  • Return on Capital Employed (ROCE)unmatched
  • Software Designunmatched
  • Software Engineeringunmatched
  • Standards Developmentunmatched
  • Startupunmatched
  • Systems Administration/Managementunmatched
  • Technical Recruitingunmatched
  • Technical Supportunmatched
  • Testingunmatched
  • Validation Testingunmatched
  • Vehicle Fleetsunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder