Senior Solutions Engineer, AI Infrastructure

VAST Data
  • Philadelphia, PA
    Today

    Job Description

    Were looking for a deeply technical Solutions Architect to help customers design, evaluate, and deploy infrastructure for large-scale AI, HPC, analytics, and data-intensive workloads.

    This is a customer-facing technical role for someone who has lived inside production infrastructure. You may have been a platform engineer, infrastructure engineer, SRE, MLOps engineer, AI infrastructure engineer, storage engineer, cloud engineer, or HPC systems engineer. What matters most is that you have built, operated, or architected real systems, and can bring that credibility into customer conversations.

    Our customers are building infrastructure at serious scale: GPU clusters, high-performance storage systems, Kubernetes platforms, distributed training environments, inference platforms, data pipelines, lakehouses, and large enterprise systems. Youll help them reason about architectures involving 10,000+ GPUs, 100PB+ of storage, high-performance networking, distributed filesystems, orchestration layers, and demanding production workloads.

    Youll own technical discovery, architecture design, PoC planning, competitive positioning, and customer technical strategy. Youll work from the first whiteboard session through evaluation, deployment planning, and production success. Youll also partner closely with product and engineering teams to bring field feedback into the roadmap.

    Were looking for someone who can go deep technically, communicate clearly, operate without a rigid playbook, and translate complex infrastructure into customer outcomes.

    Responsibilities

    • Lead technical discovery with customers across infrastructure, platform, ML, data, and executive stakeholders.
    • Design architectures for large-scale AI, HPC, analytics, and enterprise data workloads.
    • Help customers evaluate infrastructure involving GPUs, storage, networking, orchestration, and data movement.
    • Translate complex technical requirements into clear solution designs, reference architectures, and deployment guidance.
    • Debug customer issues across Linux, storage, networking, Kubernetes, schedulers, GPUs, and application workloads.
    • Build technical assets, runbooks, and field guidance for repeatable customer engagements.
    • Partner with product and engineering to communicate customer requirements, gaps, and roadmap opportunities.
    • Help customers move from architecture design to production deployment.

    Requirements

    • 8 to 12+ years of technical experience, with significant hands-on infrastructure experience.
    • Experience building, operating, or architecting production platform infrastructure.
    • Strong understanding of Linux kernel implementation details, distributed systems including PAXOS and raft, storage implementations details like NAND or write amplification, networking store/forward, load balancing designs, and production operations.
    • Experience with one or more of: GPU infrastructure, large scale HPC systems, Kubernetes platforms from scratch, MLOps, storage systems, cloud infrastructure, data platforms, or large-scale enterprise infrastructure.
    • Ability to communicate credibly with engineers, architects, technical executives, and business stakeholders.
    • Strong discovery, problem-solving, and systems debugging skills.
    • Comfort operating in ambiguous, fast-moving environments.
    • Interest in customer-facing technical work, solution design, and business outcomes.

    Preferred Experience

    • Experience with large-scale GPU clusters, distributed training, inference infrastructure, or AI platforms.
    • Experience with petabyte-scale storage or high-performance data systems.
    • Experience with Kubernetes, Slurm, Ray, Spark, or other orchestration / scheduling systems.
    • Domain Expertise with one or more of these - Lustre, Ceph, Weka, BeeGFS, GPFS, VAST, object storage, or distributed filesystems.
    • Experience with large-scale InfiniBand, RoCE, RDMA, high-performance Ethernet, or NVIDIA/Mellanox networking.
    • Direct Experience with CUDA, NCCL, DCGM, GPUDirect, checkpointing, dataset staging, or model-serving infrastructure.
    • Experience across multiple industries or customer environments.

    Numbers & Facts

    LocationPhiladelphia, PA

    Skills

    • Architectural Designunmatched
    • Artificial Intelligence (AI)unmatched
    • CUDA (Compute Unified Device Architecture)unmatched
    • Cloud Storageunmatched
    • Communication Skillsunmatched
    • Customer Relationsunmatched
    • Customer Support/Serviceunmatched
    • Data Managementunmatched
    • Debugging Skillsunmatched
    • Distributed Computingunmatched
    • Ethernetunmatched
    • Field Trialsunmatched
    • File Systemsunmatched
    • GPU (Graphics Processing Unit)unmatched
    • Large-Scale Systemsunmatched
    • Linux Kernelunmatched
    • Linux Operating Systemunmatched
    • Load Balancingunmatched
    • Multiplatform/Cross-Platformunmatched
    • NFS (Network File System)unmatched
    • Problem Solving Skillsunmatched
    • Product Engineeringunmatched
    • Product Positioningunmatched
    • Return on Capital Employed (ROCE)unmatched
    • Software Engineeringunmatched
    • Storage Softwareunmatched
    • System Architectureunmatched
    • Systems Engineeringunmatched
    • Technical Leadershipunmatched
    • Technical Strategyunmatched
    • Training Data Setsunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder