Founding Product Infrastructure Engineer EmanateFounding Product Infrastructure EngineerSan Francisco, CaliforniaAs a Founding Product Infrastructure Engineer at Emanate, you’ll build the core systems that power our AI revenue engine for companies that are building the backbone of the physical economy.
NewStaff Cluster Infrastructure Engineer ATOMSStaff Cluster Infrastructure EngineerSan Francisco, California$224,000–$284,000 / yearOur systems are designed to understand, predict, and control the real world with precision, turning complex physical operations into something more reliable, more scalable, and more productive. Our work only matters if it serves others, and we know that meaningful progress depends on the trust of the people we serve and the strength of our team — so we invest in both, creating an environment where you can do your best work and grow.
General Interest - Experience w/ AI Infrastructure Andromeda ClusterGeneral Interest - Experience w/ AI InfrastructureSan Francisco, CaliforniaIf you have firsthand experience working in AI infrastructure but don't see a role currently posted that you are qualified for, feel free to submit your resume here. If there is a future opportunity that we think aligns with your background, we will reach out.
Software Engineer, Infrastructure PylonSoftware Engineer, InfrastructureSan Francisco, CaliforniaUnlike platforms built before the AI era, Pylon enriches every interaction with deep account-level context, automates the low impact customer work, and surfaces answers before your team even has to ask. For leaders scaling AI-native support teams, Pylon lets humans and agents collaborate on customer work – investigating, resolving, and acting on every signal across every channel that matters.
AI Infrastructure Engineer, Sandbox Platform Scale AI, Inc.AI Infrastructure Engineer, Sandbox PlatformSan Francisco, CA$180,000–$225,000 / yearAs a Software Engineer on the AI Infrastructure team, you'll help build and evolve our agent sandboxing platform - the secure, high-performance code execution layer powering our agentic workflows, deployed across both internal and customer-managed environments. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training.
Machine Learning Infrastructure Engineer, Safeguards Research AnthropicMachine Learning Infrastructure Engineer, Safeguards ResearchSan Francisco, CAThis research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work.
AI Infrastructure Engineer, Model Serving Platform Scale AI, Inc.AI Infrastructure Engineer, Model Serving PlatformSan Francisco, CA$180,000–$225,000 / yearThe range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Our products provide the high-quality data and full-stack technologies that power the world's leading models, and help enterprises and governments build, deploy, and oversee AI applications that deliver real impact.
Lead Member Of Technical Staff, Inference Infrastructure CohereLead Member Of Technical Staff, Inference InfrastructureSan Francisco, CAIn this role, you will provide technical leadership across multiple teams, driving the architecture and strategy for deploying optimized NLP models to production in low latency, high throughput, and high availability environments. You will serve as a key point of contact for customers, leading the design of customized deployments to meet their specific needs, and mentoring engineers to raise the technical bar across the team.
AI Infrastructure Operations, Demand Planning AnthropicAI Infrastructure Operations, Demand PlanningSan Francisco, CARun a portfolio of bring-ups in parallel - new cloud regions, on-prem sites, neocloud blocks - with one integrated schedule spanning provider milestones, cluster creation, network turn-up, storage readiness, health burn-in, and first-workload landing. Capacity Engineering owns the data, tooling, and systems that let Anthropic plan, measure, and maximize utilization of that fleet: we partner on supply deals, wire telemetry from day zero, own the canonical capacity data layer, and build the planning and enforcement tools every research and product team relies on.
Infrastructure Capacity Planner, Demand Planning AnthropicInfrastructure Capacity Planner, Demand PlanningSan Francisco, CABuild and own the medium-range multi-resource demand forecast: accelerators by chip/interconnect class, CPU by shape, storage by tier and access pattern, egress by path, managed services by SKU - driven by model roadmap, RL/inference growth, eval volume, and retention policy rather than trend lines. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences.
AI Operations Engineer - IT/Internal Infrastructure LambdaAI Operations Engineer - IT/Internal InfrastructureSan Francisco, CaliforniaOur investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove. *Note: This position requires presence in our San Francisco of San Jose office location 4 days per week; Lambda's designated work from home day is currently Tuesday.
NewStaff Software Engineer (Cloud Infrastructure) Crusoe EnergyStaff Software Engineer (Cloud Infrastructure)San Francisco, CA$215,000–$260,000 / yearWe're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services. The ideal candidate will be hands-on with GPU rack-level troubleshooting and work closely with data center operations, engineering, and vendors to support cutting-edge infrastructure featuring the latest NVIDIA and AMD GPUs.
Senior Staff Software Engineer, Identity Infrastructure Engineering OpenAISenior Staff Software Engineer, Identity Infrastructure EngineeringSan Francisco, CaliforniaFor unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. The ideal candidate has experience building and operating large-scale, mission-critical systems with strong reliability and security requirements, and is comfortable writing production code, designing distributed systems, and driving ambiguous projects from 0 to 1 while building the operational rigor needed to run critical infrastructure over time.
Systems Generalist, GPT Infrastructure OpenAISystems Generalist, GPT InfrastructureSan Francisco, CaliforniaFor unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. Given a workload, target hardware profile, compiler and runtime context, and a trusted verifier, the system runs durable optimization campaigns that generate, compile, execute, grade, and improve candidate kernels, runtime configurations, and serving-stack changes.
Infrastructure Software Engineer, Energy Storage Redwood MaterialsInfrastructure Software Engineer, Energy StorageSan Francisco, CA$180,000–$237,500 / yearThe position partners closely with cross-functional engineering teams to translate early deployment learnings into platform improvements and drive resolution of scalability, reliability, and security challenges. This role serves as a force-multiplier across the organization, owning edge fleet management, server provisioning, and deployment automation while ensuring systems are secure, scalable, and performant.
Senior Software Engineer, Infrastructure (Sre & Security Focused) NimbleSenior Software Engineer, Infrastructure (Sre & Security Focused)San Francisco, CA$200,000–$260,000 / yearOur founding team comes from the AI labs at Stanford and Carnegie Mellon and our board of directors include famed robotics and AI legends including Fei-Fei Li (Chief Scientist of AI at Google and Director of Stanford's AI Lab), Marc Raibert (founder of Boston Dynamics), and Sebastian Thrun (founder of GoogleX, Waymo; Stanford Professor and considered the father of autonomous vehicles). We are looking for a Senior Software Engineer to join our Infrastructure Team, focused on building secure, reliable, and scalable infrastructure for Nimble's software services, robotics products, embedded systems, and business operations.
Software Engineer, Fleet Infrastructure OpenAISoftware Engineer, Fleet InfrastructureSan Francisco, CaliforniaFor unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment.
Staff Machine Learning Infrastructure Engineer ATOMSStaff Machine Learning Infrastructure EngineerSan Francisco, CaliforniaYou will own the challenge of scaling distributed GPU workloads to support a high volume of concurrent training runs across an expanding vehicle fleet, building a platform that can flexibly run on whatever GPU capacity is available, regardless of provider or environment, directly accelerating innovation across the platform. Distributed Computing & Orchestration: Leverage distributed compute frameworks to efficiently manage and execute a high volume of complex ML training jobs concurrently across large GPU clusters.
Software Engineer, Data Infrastructure - Research OpenAISoftware Engineer, Data Infrastructure - ResearchSan Francisco, CaliforniaFor unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. Collaborate with teammates to integrate datasets seamlessly into training and inference pipelines, ensuring smooth adoption and a great user experience.
Member of Technical Staff - AI Cloud Infrastructure Emerald AIMember of Technical Staff - AI Cloud InfrastructureOakland, CaliforniaThey have built or served as a core early engineer on a managed cloud or AI platform, whether at a GPU cloud, an internal machine learning platform run at scale, a hyperscaler AI service, or a HPC research computing center operated as a service. Production experience deploying or operating Lustre or a comparable parallel filesystem such as GPFS, Weka, VAST, or BeeGFS, with a solid understanding of parallel filesystem architecture, tuning, and failure modes.