Software Engineer, Caching Infrastructure OpenAISoftware Engineer, Caching InfrastructureSan Francisco, CaliforniaFor unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. We aim to provide a high-availability, multi-tenant cache platform that scales automatically with workload, minimizes tail latency, and supports a diverse range of use cases.
Infrastructure Engineer architect /Full time role HARAMAIN SYSTEMS INC.Infrastructure Engineer architect /Full time rolePalo Alto, CA$150,000–$180,000 / yearMust: Experience with any of the following: caching, cloud services, scaling systems, integrations, event-driven systems. Nice to: Experience with microservices, GraphQL APIs, or cross-team platform work.
Simulation Infrastructure Engineer OpenAISimulation Infrastructure EngineerSan Francisco, CaliforniaFor unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. Build scheduling, batching and orchestration for running very large numbers of concurrent rollouts (target tens of thousands of rollouts / large RL workloads), solve engine-level scaling (parallelization, batching multiple runs per engine), and optimize cloud/GPU runtime reliability.
NewEngineering Manager, GPU Infrastructure CohereEngineering Manager, GPU InfrastructureSan Francisco, CaliforniaA background running large Kubernetes compute fleets in production, including in multi-cloud environments: multi-cluster operations, scheduling, node health at scale, and familiarity with IaC and infrastructure monitoring. You’ve gone deep in one of the layers that make a GPU training fleet work, whether that’s cluster-wide operations, GPU networking, or hardware, and you’re willing to get hands-on and learn the rest.
Software Engineer, Fleet Infrastructure OpenAISoftware Engineer, Fleet InfrastructureSan Francisco, CaliforniaFor unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment.
Staff Product Manager, Cloud Infrastructure (Networking) Crusoe EnergyStaff Product Manager, Cloud Infrastructure (Networking)San Francisco, CA$215,000–$260,000 / yearWe're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services. Product strategy: Define the long-term vision for Crusoe's networking offerings, including cluster networking, SDN, and RDMA fabrics, to keep pace with the needs of AI training and inference at scale.
Software Engineer, Data Infrastructure - Research OpenAISoftware Engineer, Data Infrastructure - ResearchSan Francisco, CaliforniaFor unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. Collaborate with teammates to integrate datasets seamlessly into training and inference pipelines, ensuring smooth adoption and a great user experience.
Estimator, Infrastructure WebcorEstimator, InfrastructureSan Francisco, CaliforniaWorking knowledge of infrastructure construction costs, production-based estimating, and unit-pricing contract structures; Water/Wastewater Treatment Plants and Pump Stations, Substations and Electrical Conveyance, Mass Transit Stations, Horizontal utility development and distribution at a campus level. Prior field experience in civil construction supporting infrastructure work is preferred; Water/Wastewater Treatment Plants and Pump Stations, Substations and Electrical Conveyance, Mass Transit Stations, Horizontal utility development and distribution at a campus level.
Software Engineer, GPU Infrastructure - HPC OpenAISoftware Engineer, GPU Infrastructure - HPCSan Francisco, CaliforniaFor unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. Experience with low level details of hardware components, protocols, and associated Linux tooling (e.g., PCIe, Infiniband, networking, power management, kernel perf tuning).
Staff Software Development Engineer - Enterprise AI Infrastructure - #4898 GRAILStaff Software Development Engineer - Enterprise AI Infrastructure - #4898Menlo Park, CA$169,000–$224,000 / yearThe Staff Engineer partners closely with cross-functional stakeholders across Software Engineering, Data Science, Security, Regulatory Affairs, and Product to build a unified control plane that securely connects large language models with enterprise tools and company knowledge. We have built a multi-disciplinary organization of scientists, engineers, and physicians and we are using the power of next-generation sequencing (NGS), population-scale clinical studies, and state-of-the-art computer science and data science to overcome one of medicine’s greatest challenges.
Senior AI Infrastructure Engineer - Model Training KodiakSenior AI Infrastructure Engineer - Model TrainingMountain View, California$190,000–$260,000 / yearShould the position require, and Kodiak determines that a candidate’s residence, U.S. person status, and/or citizenship status necessitate an export license, bar the candidate from the position, or otherwise fall under national security-related restrictions, Kodiak will consider the candidate for alternative positions unaffected by such restrictions, under terms and conditions set forth at Kodiak’s sole discretion, or, as an alternative, opt not to proceed with the candidate’s application. Experience building high-performance data pipelines for large-scale training, including streaming dataset formats (WebDataset, MosaicML Streaming/MDS, or similar), sharding, and storage/network-aware loading.
NewInfrastructure Design Engineer Crusoe EnergyInfrastructure Design EngineerSunnyvale, CA$155,000–$190,000 / yearWe're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services. The Infrastructure Design Engineer is responsible for planning, designing, and implementing data center whitespace environments - the operational computing areas where servers, storage, and network equipment are deployed.
Mechanical Systems Inspector - Water Infrastructure AtlasMechanical Systems Inspector - Water InfrastructureOakland, CA$55–$65 / hourWe trust our people to do great work and live fulfilling lives, and we provide resources that care for our people, both at work and in life: Health & Wellness: We offer medical, dental and vision coverage with multiple plan options, including no-cost employee-only medical coverage. In the event a recruiter or agency submits a resume or candidate without a previously signed agreement, Atlas explicitly reserves the right to pursue and hire those candidate(s) without any financial obligation to the recruiter or agency.
AI Training Infrastructure Engineer – Humanoid Whole Body Control FigureAI Training Infrastructure Engineer – Humanoid Whole Body ControlSan Jose, California$150,000–$300,000 / yearThis role sits at the intersection of robotics, machine learning, controls, and software systems engineering, and is critical to how quickly we can iterate, train, and deploy new capability to our fleet of humanoid robots. Experience building or scaling training infrastructure for robotics, control systems, or large-scale ML workloads.
Lightning Infrastructure Engineer (Remote) Lightning LabsLightning Infrastructure Engineer (Remote)Palo Alto, CaliforniaRemoteIn addition to the core Lightning Network Daemon software and end-user applications, the Lightning Network ecosystem includes supporting systems like watchtowers (a form of backup), peer availability/network monitoring, advanced liquidity provisioning tools, automated channel management, and other services. These tools will lower the barrier to entry for operating routing nodes[1] and enable existing routing node operators to more effectively manage their infrastructure.
Global Public Policy Manager, Compute, Infrastructure & Sovereign AI CohereGlobal Public Policy Manager, Compute, Infrastructure & Sovereign AISan Francisco, California$190,000–$230,000 / yearAll legitimate roles are listed on the Cohere careers page and LinkedIn only, with all communications from Cohere employees coming from an @cohere.com or @cw.cohere email alias. Governments increasingly view AI as critical national infrastructure and are investing heavily in compute capacity, energy resources, and domestic AI ecosystems.
Software Engineering Manager, Platform Pure Storage Inc.Software Engineering Manager, PlatformSanta Clara, CA$217,000–$326,000 / yearClassification of protected categories is as follows: A "disabled veteran" is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service-connected disability. In this role, you will accelerate the execution of core platform software integrating cutting-edge Linux kernel, networking, and NVMe technologies, collaborating directly with Product and Release management to drive massive business impact in the AI neocloud market.
Software Engineer, Infrastructure - Analytics Platform OpenAISoftware Engineer, Infrastructure - Analytics PlatformSan Francisco, CaliforniaFor unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. This person should be able to own a platform such as CacheHouse end to end: define its technical direction, design its data model and storage architecture, integrate it with several research dashboards and workflows, guide one or two engineers, and ensure the system works reliably for its users.
AI Training Infrastructure Engineer - Humanoid Whole Body Control FigureAI Training Infrastructure Engineer - Humanoid Whole Body ControlSan Jose, CA$150,000–$300,000 / yearThis role sits at the intersection of robotics, machine learning, controls, and software systems engineering, and is critical to how quickly we can iterate, train, and deploy new capability to our fleet of humanoid robots. Experience building or scaling training infrastructure for robotics, control systems, or large-scale ML workloads.
Technical Program Manager, Infrastructure AnthropicTechnical Program Manager, InfrastructureSan Francisco, CAEvery breakthrough in AI safety research and every interaction users have with Claude depends on the systems we build and operate: massive clusters for training frontier models, production infrastructure serving millions of users reliably, and developer platforms that help engineers move fast without breaking things. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences.