Principal Deployment Engineer

Nscale AS

  • Seattle, WA
  • 30+ days ago
    Want to know if you’re a fit?
    Upload your resume and let our AI show you.

    Skills

    • Artificial Intelligence (AI)unmatched
    • Automationunmatched
    • Benchmarkingunmatched
    • Best Practicesunmatched
    • Blueprintsunmatched
    • Broadbandunmatched
    • Capacity and Performance Managementunmatched
    • Computer Firmwareunmatched
    • Computer Networksunmatched
    • Cross-Functionalunmatched
    • Debugging Skillsunmatched
    • Distributed Computingunmatched
    • Ethernetunmatched
    • GPU (Graphics Processing Unit)unmatched
    • Leadershipunmatched
    • Memory Hardwareunmatched
    • Mentoringunmatched
    • Network Operations Centerunmatched
    • PCI Express (PCI-E)unmatched
    • Return on Capital Employed (ROCE)unmatched
    • Server Architectureunmatched
    • Service Level Agreement (SLA)unmatched
    • Standards Developmentunmatched
    • Startupunmatched
    • Technical Leadershipunmatched
    • Time Managementunmatched
    • Topologyunmatched

    Description

    Principal Deployment Engineer - GPU Supercluster Bringup

    About Us We are building AI infrastructure for frontier-scale workloads. Our platform is designed for high-density, high-performance GPU clusters that push the limits of power, networking, and distributed compute.

    As a startup, we move fast, operate with ownership, and expect technical leaders to define standards-not just follow them.

    The Role We are hiring a Principal Deployment Engineer to architect and lead the bringup of large-scale GPU clusters (hundreds to thousands of GPUs). This is a technical leadership role responsible for defining how we deploy, validate, and scale AI superclusters across sites.

    You will own the full lifecycle of deployment-from rack design and fabric architecture to cluster validation frameworks and production readiness standards. You will set the bar for performance, reliability, and operational excellence.

    This role combines deep hands-on expertise with system-level thinking and cross-functional leadership.

    What You'll Do

    End-to-End Supercluster Bringup Ownership

    • Define the technical standards for node, rack, and full-cluster bringup.

    Lead large-scale GPU cluster deployments (multi-rack, multi-pod environments).

    • Architect high-performance network fabrics (IB, RoCE, Ethernet) optimized for AI workloads.
    • Establish cluster-level acceptance criteria and validation frameworks.

    Performance & Fabric Architecture

    • Tune and validate NCCL, RDMA, GPUDirect, and collective operations at scale.
    • Identify and eliminate performance bottlenecks across hardware, topology, and firmware layers.
    • Drive congestion control and fabric optimization strategies.
    • Define performance benchmarking methodology for AI training workloads.

    Deployment Strategy & Scalability

    • Design repeatable deployment models for multi-site expansion.
    • Build automation frameworks for provisioning and cluster validation.
    • Establish deployment SLAs, quality gates, and operational readiness standards.
    • Reduce time-to-capacity while increasing reliability.

    Technical Leadership

    • Serve as the escalation point for complex bringup and performance issues.
    • Mentor senior engineers and shape infrastructure best practices.
    • Influence hardware selection, rack topology, and data center design decisions.
    • Partner with executive leadership on infrastructure scaling strategy.

    What We're Looking For

    Required

    • 10+ years of experience in large-scale infrastructure or HPC environments.
    • Proven experience bringing up large GPU clusters (hundreds+ GPUs).
    • Deep expertise in high-speed networking (InfiniBand, RoCE, Ethernet fabrics).
    • Strong understanding of server architecture (PCIe, NUMA, memory hierarchy).
    • Experience debugging performance issues across compute and network layers.
    • Strong automation and systems-level thinking.

    Strongly Preferred

    • Experience scaling AI training clusters for frontier models.
    • Experience with liquid cooling or ultra-high-density deployments.
    • Knowledge of distributed storage systems (Lustre, Ceph, NVMe-oF).
    • Experience defining infrastructure standards in a fast-growing organization.

    What Success Looks Like

    • Superclusters are brought online quickly, predictably, and at peak performance.
    • Deployment processes scale from first cluster to multi-site expansion.
    • Infrastructure becomes a competitive advantage.
    • You define the technical blueprint for how we scale AI infrastructure.

    Numbers & Facts

    LocationSeattle, WA

    Similar Jobs

    See more jobs