Neo Cloud - Principal AI Cloud Storage Engineer

Blaze Talent
  • San Francisco, California
    2 days ago

    Job Description

    Member of Technical Staff, AI Cloud Storage

    Location: Bay Area/ Seattle/ Remote
    Reports to: CTO

    About Neo Cloud
    Neo Cloud is a fast growing next-generation AI neocloud, founded by ex NVIDIA, CoreWeave and Intel engineering leaders. We offer you the opportunity to build and operate next generation AI infrastructure that powers the world's most demanding AI workloads for the largest AI labs. We are rethinking how to run large scale AI infra for instance by leveraging agentic AI to build digital twins to de-risk our physical deployments. Our agents are deploying and monitoring our fleet to maximize the uptime of our infra.

    About the Role
    We're looking for a Principal Software Engineer to define and build Neo Cloud's AI cloud storage platform. This is a senior individual-contributor role for an experienced system architect who can operate at the intersection of customer needs, system architecture, and production operations — translating what AI/ML customers will need in the future, into a storage system that is performant, resilient, and economical at massive scale.

    Key Responsibilities
    Identifying customer requirements:
    • Engage directly with customers, solutions architects, and product teams to understand future storage requirements for AI/ML workloads.
    • Extrapolate from current usage patterns and industry trends to anticipate future requirements.
    • Partner with product management to prioritize platform investments based on near-term customer needs and longer-term strategic bets.

    System design, implementation, and operations:
    • Design, implement, and operate AI cloud object and file storage systems, including the data path, metadata path, and control plane.
    • Take a system-level approach that accounts for the full characteristics of AI workloads, building end-to-end solutions — including local node caching strategies, the network hardware and protocols that move data between storage and compute (e.g., RDMA, high-throughput NICs, congestion control), and the underlying storage hardware and software (media, erasure coding, metadata services) — and understanding how decisions in one layer constrain or unlock the others.
    • Architect for the specific demands of AI workloads: very high aggregate throughput to keep GPU/accelerator clusters fed, support for massive numbers of small and large objects and files, efficient checkpointing at scale, and predictable tail latency under heavy concurrent load.
    • Drive core storage system design decisions, including durability and consistency models, erasure coding and replication strategies, metadata scalability, multi-tenancy and isolation, and S3-compatible and POSIX/file-protocol API design.
    • Take end-to-end ownership of services in production: build for observability and operability from day one, participate in on-call, lead incident response and root-cause analysis for critical issues, and drive long-term reliability and performance improvements.
    • Identify and eliminate performance bottlenecks and scalability limits before they become customer-facing problems; lead capacity planning for rapid growth.
    • Deep understanding of data privacy and security and its implication on performance.
    • Ability to partner with network engineers to deliver complete AI storage system.

    Technical leadership:
    • Strong bias to action and resolution of technical decisions. 
    • Set technical direction and best practices for the storage organization; author and review design documents for significant architectural changes.
    • Provide deep technical mentorship to senior and staff engineers; raise the engineering bar across the team through code review, design review, and hands-on collaboration.
    • Influence technical strategy across adjacent teams.

    Qualifications
    • 10+ years of professional software engineering experience building and operating cloud storage systems in production.
    • Direct experience with AI-focused storage platforms such as DDN, Weka, or VAST, including their architectural approaches to throughput, caching, and GPU-cluster integration.
    • Deep understanding of distributed systems fundamentals: consistency models, replication, consensus, partitioning, failure detection, and recovery.
    • Proven experience operating high-scale distributed systems in production, including on-call ownership, incident response, and driving systemic reliability improvements.
    • Strong systems programming skills (e.g., Go, C++, Rust, or Java) and comfort working across the stack from low-level I/O and networking to distributed control planes.
    • Excellent written and verbal communication skills.

    Nice to have
    • Experience designing or tuning local node caching layers to accelerate AI training and inference data access.
    • Experience with high-performance networking.
    • Contributions to open-source storage projects, relevant patents, or published technical papers/talks.




    Numbers & Facts

    LocationSan Francisco, California

    Skills

    • Amazon Simple Storage Service (S3)unmatched
    • Application Programming Interface (API)unmatched
    • Architectural Designunmatched
    • Architectural Servicesunmatched
    • Artificial Intelligence (AI)unmatched
    • Best Practicesunmatched
    • C++ Programming Languageunmatched
    • Cachingunmatched
    • Capacity Managementunmatched
    • Cloud Computingunmatched
    • Cloud Storageunmatched
    • Code Reviewsunmatched
    • Communication Skillsunmatched
    • Computer Programmingunmatched
    • Computer Storage Hardwareunmatched
    • Customer Relationsunmatched
    • Data Storageunmatched
    • Design Documentunmatched
    • Distributed Computingunmatched
    • Engineeringunmatched
    • Establish Prioritiesunmatched
    • File Systemsunmatched
    • GPU (Graphics Processing Unit)unmatched
    • Go Programming Language (Golang)unmatched
    • High Throughputunmatched
    • Incident Responseunmatched
    • Industry/Trade Analysisunmatched
    • Information/Data Security (InfoSec)unmatched
    • Input/Outputunmatched
    • Intel Product Familyunmatched
    • Javaunmatched
    • Mentoringunmatched
    • Metadataunmatched
    • Network Architecture/Engineeringunmatched
    • Network Protocolsunmatched
    • Network System Hardwareunmatched
    • On Callunmatched
    • Open Sourceunmatched
    • POSIX Operating Systemsunmatched
    • Patentsunmatched
    • Performance Managementunmatched
    • Presentation/Verbal Skillsunmatched
    • Privacy Controlsunmatched
    • Process Improvementunmatched
    • Product Managementunmatched
    • Production Systemsunmatched
    • Protocol Designunmatched
    • Publicationsunmatched
    • Reliability Engineeringunmatched
    • Replication and Remote Mirroringunmatched
    • Riskunmatched
    • Root Cause Analysisunmatched
    • Rust Programming Languageunmatched
    • Software Engineeringunmatched
    • Storage Softwareunmatched
    • System Architectureunmatched
    • Systems/Internals Programmingunmatched
    • Technical Leadershipunmatched
    • Technical Publicationsunmatched
    • Technical Strategyunmatched
    • Vehicle Fleetsunmatched
    • Writing Skillsunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder