Nscale is the vertically integrated AI cloud engineered for AI. We own and operate the full stack — energy, data centres, GPU superclusters, orchestration, and AI services — delivering high-performance infrastructure to AI-native companies, enterprises, and governments across Europe and the US. We are deploying GPU capacity at hyperscale, operating some of the densest, most advanced AI infrastructure in the world.
At Nscale, our Support and Operations team plays a critical role in maintaining service availability, driving service reliability, and delivering rapid response to customer issues. We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.
The Head of Infrastructure Support owns Infrastructure Support for their region — the team, the function, and its impact on customers. Reporting directly to the VP of Support and operating alongside counterpart Heads of Infrastructure Support across EMEA, the US, and APAC, you are accountable for the success of regional support outcomes: service performance, escalation quality, customer experience, and the health of the GPU estates your team supports.
The regional Infrastructure Support engineers report directly to you, and you own their management end to end — hiring, 1:1s, performance reviews, development planning, and documented performance management through to outcome. Their performance, growth, and results are your responsibility.
You will grow your regional Infrastructure Support team during a period of rapid company scaling, embed a consistent operating model with your counterparts in the other regions to deliver true follow-the-sun coverage, and act as the organisational accountability layer for your region — ensuring that strategic and tactical work spanning Support and Operations lands with clear owners and gets driven to completion. As the function matures, you will own a global capability area on behalf of all regions and shape your team's structure — including developing team leads — as headcount grows.
You remain technically credible: close enough to GPU infrastructure, high-performance fabrics, and Linux operations to lead complex incident response, challenge technical decisions on their merits, and earn the respect of Senior Engineers — while spending the majority of your time leading.
Experience required:
8+ years in infrastructure, operations, or support engineering in production environments, including 4+ years of direct line management of engineers in an operational support function, with demonstrable ownership of performance management. Significant exposure to GPU, HPC, or large-scale data centre estates.
5+ years of direct line management of engineers in an operational support environment, with end-to-end ownership of performance management: reviews, development plans, and documented underperformance processes through to outcome. You can describe your management framework and point to engineers you've grown.
Experience owning team workload, prioritisation, and service delivery against SLAs, with accountability for the numbers — and experience explaining those numbers to senior leadership.
Experience hiring, scaling, or standing up support/operations capability in a fast-moving environment, including capacity modelling and headcount planning against demand; comfortable operating where processes are still evolving and helping define them without slowing delivery.
Excellent written and verbal communication — clear, specific, and concise at every level, from ticket notes to executive updates to difficult customer conversations. We treat communication quality as a core leadership skill and assess it directly.
A bias for decisive action and calculated risk in ambiguous situations; you take ownership of outcomes, speak candidly, disagree when appropriate, and commit fully once decisions are made.
Strong understanding of ITIL-aligned incident, problem, and change management, and of SRE practices — runbooks, toil reduction, and continuous improvement.
Comfortable with out-of-hours escalations, regional on-call participation, and travel for onsite leadership.
NCCL-based performance troubleshooting, NVLink/NVSwitch, Slurm-scheduled multi-GPU workloads, or rack-scale systems.
Exposure to VAST or comparable AI-optimised storage platforms, Ceph, or NFS at scale.
OpenStack operations, or fleet-scale provisioning and health tooling (MAAS, NetBox, Redfish-driven automation, or similar).
Operating or supporting clusters and GPU operator stacks. Helpful context for our platform, though not the core of this role.
Experience running or coordinating support across regions and time zones.
ITIL certification, or relevant Linux/networking/cloud certifications.
The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.
Salary Range
$200,000 - $300,000USD
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.
| Location | Houston, Texas |
| Website | https://www.nscale.com |
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.
Free resume templatesImprove your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.
Free resume builder