About the team:
The Infra-Compute division builds large-scale, highly available cloud and AI infrastructure that powers our public cloud offerings and internal products. Our US team develops technologies across AI training, inference, and agent infrastructure.
We are expanding this work into Physical AI: intelligent systems that perceive, reason, and act in the physical world. The team focuses on infrastructure for Physical AI, including data and simulation platforms, distributed training and inference, and deployment on heterogeneous hardware. We collaborate closely with customers, researchers, open-source communities, and hardware partners to turn emerging research into reliable, scalable systems.
Responsibilities:
Independently own the design, implementation, validation, and operation of major components within a Physical AI platform, focusing on one or more areas across data pipelines, simulation, model training, evaluation, inference, and heterogeneous deployment.
Develop scalable infrastructure for training and serving multimodal and robotics models, including vision-language-action models, world models, and robot policies.
Translate research ideas into production-ready capabilities, with clear reliability, performance, and quality targets.
Lead technical design for complex projects, make well-reasoned trade-offs across performance, reliability, cost, and delivery speed, and drive projects through production adoption.
Collaborate with adjacent teams, mentor engineers, and contribute to shared architecture and engineering standards.Minimum Qualifications:
Bachelor's degree or above in Computer Science, Computer Engineering, Robotics, Electrical Engineering, Artificial Intelligence, or a related technical field, or equivalent practical experience.
5 years of software engineering, machine learning systems, robotics, or related industry experience.
Strong software and systems engineering fundamentals, with experience designing, building, and operating production-scale distributed or data-intensive systems.
Deep expertise in at least one of the following areas:
Robotic learning, including imitation learning, reinforcement learning, vision-language-action models, world models, or policy evaluation.
AI training and inference infrastructure, including distributed training, inference engines, GPU kernels, or collective communication.
Robotics data and simulation systems, including multimodal dataset verification, synthetic data generation, simulation, data curation, or sim-to-real workflows.
Preferred Qualifications
Contributions to relevant open-source projects such as Isaac Lab, Isaac Sim, MuJoCo, Cosmos, vLLM-Omni, SGLang Omni, RLInf, etc.
Experience building large-scale AI or robotics platforms used by multiple teams or external customers, or experience with real-world robot systems and operational challenges.
Numbers & Facts
Location
Seattle, WA
Skills
Artificial Intelligence (AI)unmatched
Cloud Computingunmatched
Computer Engineeringunmatched
Computer Scienceunmatched
Data Managementunmatched
Data Setsunmatched
Electrical Engineeringunmatched
GPU (Graphics Processing Unit)unmatched
High Availabilityunmatched
Inference Engineunmatched
Kernel Programmingunmatched
Machine Learningunmatched
Mentoringunmatched
Modeling Languagesunmatched
Open Sourceunmatched
Policy Evaluationunmatched
Public Cloudunmatched
Reinforcement Learningunmatched
Roboticsunmatched
Scalable System Developmentunmatched
Simulationunmatched
Software Engineeringunmatched
System Operationsunmatched
Systems Engineeringunmatched
Systems Scalabilityunmatched
Technical Leadershipunmatched
Technical/Engineering Designunmatched
🎯
Be found by employers
5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.
Level up your application
Professional resume templates
Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.