Staff AI Runtime Engineer

FlexAI
  • San Jose, California
    30+ days ago

    Job Description

    Role Overview

    At FlexAI, we’re building a high-performance, cloud-agnostic AI compute platform designed for next-generation training and inference workloads. As a Staff AI Runtime Engineer, you’ll play a pivotal role in the design, development, and optimization of the core runtime infrastructure that powers distributed training and deployment of large AI models (LLMs and beyond).


    This is a hands-on leadership role - perfect for a systems-minded software engineer who thrives at the intersection of AI workloads, runtimes, and performance-critical infrastructure. You’ll own critical components of our PyTorch-based stack, lead technical direction, and collaborate across engineering, research, and product to push the boundaries of elastic, fault-tolerant, high-performance model execution.

    What You'll Do

    Lead Runtime Design & Development:

    • Own the core runtime architecture supporting AI training and inference at scale.
    • Design resilient and elastic runtime features (e.g. dynamic node scaling, job recovery) within our custom PyTorch stack.
    • Optimize distributed training reliability, orchestration, and job-level fault tolerance.

    Drive Performance at Scale:

    • Profile and enhance low-level system performance across training and inference pipelines.
    • Improve packaging, deployment, and integration of customer models in production environments.
    • Ensure consistent throughput, latency, and reliability metrics across multi-node, multi-GPU setups.

    Build Internal Tooling & Frameworks:

    • Design and maintain libraries and services that support model lifecycle: training, checkpointing, fault recovery, packaging, and deployment.
    • Implement observability hooks, diagnostics, and resilience mechanisms for deep learning workloads.
    • Champion best practices in CI/CD, testing, and software quality across the AI Runtime stack.

    Collaborate & Mentor:

    • Work cross-functionally with Research, Infrastructure, and Product teams to align runtime development with customer and platform needs.
    • Guide technical discussions, mentor junior engineers, and help scale the AI Runtime team’s capabilities.


    What You’ll Need to Be Successful

    • 8+ years of experience in systems/software engineering, with deep exposure to AI runtime, distributed systems, or compiler/runtime interaction.
    • Experience in delivering PaaS services.
    • Proven experience optimizing and scaling deep learning runtimes (e.g. PyTorch, TensorFlow, JAX) for large-scale training and/or inference.
    • Strong programming skills in Python and C++ (Go or Rust is a plus).
    • Familiarity with distributed training frameworks, low-level performance tuning, and resource orchestration.
    • Experience working with multi-GPU, multi-node, or cloud-native AI workloads.
    • Solid understanding of containerized workloads, job scheduling, and failure recovery in production environments.

    Nice to Have

    • Contributions to PyTorch internals or open-source DL infrastructure projects.
    • Familiarity with LLM training pipelines, checkpointing, or elastic training orchestration.
    • Experience with Kubernetes, Ray, TorchElastic, or custom AI job orchestrators.
    • Background in systems research, compilers, or runtime architecture for HPC or ML.
    • Start up previous experience

    This position is In-Person and located at our Santa Clara, CA Office.

    What We Offer

    • A competitive salary and benefits package
    • Work on cutting-edge AI infrastructure
    • Build products used by developers and enterprises
    • High ownership, fast execution, real impact
    • Collaborative, high-caliber team

    Numbers & Facts

    LocationSan Jose, California

    Skills

    • Artificial Intelligence (AI)unmatched
    • Best Practicesunmatched
    • C++ Programming Languageunmatched
    • Cloud Computingunmatched
    • Computer Programmingunmatched
    • Computer Systemsunmatched
    • Continuous Deployment/Deliveryunmatched
    • Continuous Integrationunmatched
    • Cross-Functionalunmatched
    • Deep Learningunmatched
    • Distributed Computingunmatched
    • GPU (Graphics Processing Unit)unmatched
    • JAX (Java API for XML)unmatched
    • Leadershipunmatched
    • Machine Toolunmatched
    • Mentoringunmatched
    • Metricsunmatched
    • Open Sourceunmatched
    • Performance Modelingunmatched
    • Performance Tuning/Optimizationunmatched
    • Platform as a Service (PaaS)unmatched
    • Product Developmentunmatched
    • Production Systemsunmatched
    • Python Programming/Scripting Languageunmatched
    • Service Deliveryunmatched
    • Software Engineeringunmatched
    • Startupunmatched
    • Team Playerunmatched
    • Technical Leadershipunmatched

    Be found by employers

    5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

    Level up your application

    Professional resume templates

    Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

    Free resume templates

    Free resume builder

    Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

    Free resume builder