Neo Cloud - Principal Hardware Systems Engineer

Blaze Talent
  • San Francisco, California
    2 days ago

    Job Description

    Member of Technical Staff, Hardware Engineering

    Location: Bay Area/ Seattle/ Remote
    Reports to: CTO

    About Neo Cloud
    Neo Cloud is a fast growing next-generation AI neocloud, founded by ex NVIDIA, CoreWeave and Intel engineering leaders. We offer you the opportunity to build and operate next generation AI infrastructure that powers the world's most demanding AI workloads for the largest AI labs. We are rethinking how to run large scale AI infra for instance by leveraging agentic AI to build digital twins to derisk our physical deployments. Our agents are deploying and monitoring our fleet to maximize the uptime of our infra.

    About the Role
    We're looking for a Principal Hardware Systems Engineer to design, deploy and manage Neo Cloud's AI hardware. This is a senior individual-contributor role for an experienced system architect who can operate at the intersection of customer needs, system architecture, and production operations — translating what AI/ML customers will need in the future, into a hardware design and operation at massive scale.

    Key Responsibilities
    • Identifying customer requirements
    • Engage directly with customers, solutions architects, and product teams to understand future hardware requirements for AI/ML workloads.
    • Extrapolate from current usage patterns and industry trends to anticipate future requirements.
    • Partner with vendors and drive their roadmaps.

    Hardware design, deployment, and operations:
    • Design, implement, and operate AI cloud hardware.
    • Architect, evaluate, and Deploy NVIDIA GPU infrastructure, including Blackwell-based B200, GB200, GB300, DGX and NVL rack-scale systems.
    • Design rack-scale infrastructure encompassing NVLink and NVSwitch fabrics, Infiniband or Spectrum-X Ethernet with RoCE for scale-out networking; power delivery, liquid cooling, cabling and serviceability.
    • Take a system-level approach that accounts for the full characteristics of AI workloads, building end-to-end solutions.
    • Architect for the specific demands of AI workloads focusing on performance, security, resiliency and cost. 
    • Design and implement consistent and efficient hardware operation. 
    • Take end-to-end ownership of hardware from design, supply chain planning, deployment, operation and ultimately deprecation. 
    • Ensure security and compliance. 
    • Strong debugging skills and analytical skills to identify failures across a variety of electrical, thermal, optical, mechanical and software failures. 

    Technical leadership:
    • Strong bias to action and resolution of technical decisions. 
    • Set technical direction and best practices for the hardware organization; author and review design documents for significant architectural changes.
    • Provide deep technical mentorship to senior and staff engineers; raise the engineering bar across the team through code review, design review, and hands-on collaboration.
    • Influence technical strategy across adjacent teams.

    Qualifications
    • 10+ years of professional hardware engineering experience building and operating hardware at massive scale.
    • Direct experience with AI-focused hardware deployments across different hardware vendors, including NVIDIA GPU platforms.
    • Experience with NVIDIA BLackwell, GB200 or GB300 NVL72.
    • Experience with NVLink, NVSwitch, InfiniBand or Spectrum-X Ethernet, ConnectX networking, and BlueField DPUs.
    • Deep system level expertise. 
    • Proven experience operating large-scale distributed systems in production, including on-call ownership, incident response, and driving systemic reliability improvements.
    • Strong systems programming skills.
    • Excellent written and verbal communication skills.


    Numbers & Facts

    LocationSan Francisco, California

    Skills

    • Analysis Skillsunmatched
    • Architectural Analysisunmatched
    • Architectural Designunmatched
    • Artificial Intelligence (AI)unmatched
    • Best Practicesunmatched
    • Cloud Computingunmatched
    • Code Reviewsunmatched
    • Communication Skillsunmatched
    • Computer Engineeringunmatched
    • Computer Programmingunmatched
    • Debugging Skillsunmatched
    • Design Documentunmatched
    • Distributed Computingunmatched
    • Electricityunmatched
    • Engineeringunmatched
    • Ethernetunmatched
    • Failure Analysisunmatched
    • GPU (Graphics Processing Unit)unmatched
    • Hardware Designunmatched
    • Incident Responseunmatched
    • Industry/Trade Analysisunmatched
    • Intel Product Familyunmatched
    • Large-Scale Systemsunmatched
    • Maintain Complianceunmatched
    • Mentoringunmatched
    • On Callunmatched
    • Presentation/Verbal Skillsunmatched
    • Process Improvementunmatched
    • Production Systemsunmatched
    • Reliability Engineeringunmatched
    • Return on Capital Employed (ROCE)unmatched
    • Supply Chain Managementunmatched
    • System Architectureunmatched
    • Systems Engineeringunmatched
    • Systems/Internals Programmingunmatched
    • Technical Leadershipunmatched
    • Technical Strategyunmatched
    • Vehicle Fleetsunmatched
    • Writing Skillsunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder