ML Infrastructure Engineer zaimlerML Infrastructure EngineerSan Mateozaimler was founded by Biswajit Das (ex-VP Engineering, Truera), a Data Infra veteran and former Chief Architect at Visa, and Sofus Macskassy (ex-Director of Engineering, LinkedIn), who built one of the largest knowledge graphs in production in the industry at LinkedIn. zaimler is the context infrastructure for the agentic era: a platform that automatically discovers domain knowledge, maps relationships, and gives AI agents the semantic understanding to operate with precision at scale.
Member of Technical Staff — Cluster Infrastructure & Supercomputing RadixArkMember of Technical Staff — Cluster Infrastructure & SupercomputingPalo Alto, CaliforniaOur team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs. You will design and operate highly reliable, high-performance GPU/TPU clusters, build next-generation scheduling and resource management systems, and push the limits of large-scale distributed infrastructure for AI workloads.
Research Member of Technical Staff- Data Infrastructure Rhoda AIResearch Member of Technical Staff- Data InfrastructureMountain View, CaliforniaOur robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We're looking for Data Infrastructure MLEs to scale the systems that power our model training data pipeline, from raw ingestion and storage to indexing, retrieval, and throughput optimization at massive scale.
Cloud Infrastructure Engineer zaimlerCloud Infrastructure EngineerSan Mateozaimler was founded by Biswajit Das (ex-VP Engineering, Truera), a Data Infra veteran and former Chief Architect at Visa, and Sofus Macskassy (ex-Director of Engineering, LinkedIn), who built one of the largest knowledge graphs in production in the industry at LinkedIn. zaimler is the context infrastructure for the agentic era: a platform that automatically discovers domain knowledge, maps relationships, and gives AI agents the semantic understanding to operate with precision at scale.
Member of Technical Staff — Reliability-CI Infrastructure RadixArkMember of Technical Staff — Reliability-CI InfrastructurePalo Alto, CaliforniaRadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang (30K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.
CAD & Engineering Infrastructure Lead AuradineCAD & Engineering Infrastructure LeadSanta Clara, CA$200,000–$500,000 / yearThis role owns flow release and qualification, design-flow CI/CD pipelines, multi-vendor EDA tool deployment, and EDA licensing, procurement, and vendor strategy; domain design methodology and flow content stay with the functional teams: RTL, verification, DFT, physical design, and packaging. We are looking for a CAD & Engineering Infrastructure Lead to build and own the engineering platform behind Velaura's silicon development: the compute infrastructure, EDA ecosystem, integrated RTL-to-GDS release flow, and AI-assisted workflows that every design team depends on.
Technical Program Manager III, NetInfra, Technical Infrastructure Program Management Office GoogleTechnical Program Manager III, NetInfra, Technical Infrastructure Program Management OfficeSunnyvale, CAAs a PLANET Networking Technical Program Manager, you will orchestrate the lifecycle of critical networking infrastructure spanning accelerator and machine learning networks, advanced switching architectures, and SmartNIC technologies ensuring seamless integration and deployment through strategic cross-functional leadership. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.
Senior Technical Program Manager I, Data Center NPI, Cloud Infrastructure GoogleSenior Technical Program Manager I, Data Center NPI, Cloud InfrastructureSunnyvale, CADefine and drive requirements for serviceability, deployability, and maintainability of new products in the data center environment while acting as the primary liaison between engineering design teams and data center operations teams. Identify potential risks to program success (e.g., technical, schedule, operational), as well as drive root cause analysis and corrective actions for issues arising during new product introduction (NPI).
Senior Software Engineering Manager, Network Test and Infrastructure GoogleSenior Software Engineering Manager, Network Test and InfrastructureSunnyvale, CATeams work all across the company, in areas such as information retrieval, artificial intelligence, natural language processing, distributed computing, large-scale system design, networking, security, data compression, user interface design; the list goes on and is growing every day. With technical and leadership expertise, you manage engineers across multiple teams and locations, a large product budget and oversee the deployment of large-scale projects across multiple sites internationally.
Software Engineering Manager II, Infrastructure, Google Cloud Networking GoogleSoftware Engineering Manager II, Infrastructure, Google Cloud NetworkingSunnyvale, CATeams work all across the company, in areas such as information retrieval, artificial intelligence, natural language processing, distributed computing, large-scale system design, networking, security, data compression, user interface design; the list goes on and is growing every day. With technical and leadership expertise, you manage engineers across multiple teams and locations, a large product budget and oversee the deployment of large-scale projects across multiple sites internationally.
Manager, Datacenter Compute Infrastructure (Starlink) Space Exploration TechnologiesManager, Datacenter Compute Infrastructure (Starlink)Palo Alto, CA$155,000–$215,000 / yearITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. This role oversees GPU/accelerators, CPUs, storage, ODM/CM relationships, as well as server and rack components to support SpaceXAI's large-scale AI training clusters and compute facilities.
Engineering Manager, AI Models Infrastructure IntercomEngineering Manager, AI Models InfrastructureDublin, CAFounded in 2011, Fin became one of the fastest growing companies and remains one of the largest private software companies in the world with nearly 30,000 global businesses using our products to transform their customer support. Fin can also be combined with our Helpdesk to become a complete solution called the Fin Customer Service Suite, which provides AI enhanced support for the more complex or high touch queries that require a human agent.
NewSr. Manager, Datacenter Compute Infrastructure (Starlink) Space Exploration TechnologiesSr. Manager, Datacenter Compute Infrastructure (Starlink)Palo Alto, CA$185,000–$260,000 / yearITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. This role oversees GPU/accelerators, CPUs, storage, ODM/CM relationships, as well as server and rack components to support SpaceXAI's large-scale AI training clusters and compute facilities.
NewTechnical Project Manager - Cloud, Infrastructure & AI Operations Infobahn SoftWorld IncTechnical Project Manager - Cloud, Infrastructure & AI OperationsSanta Clara, CA$65–$70The ideal candidate is a strategic and execution-focused technical program leader who can operate effectively in a complex enterprise environment, translate technical strategy into executable programs, manage competing priorities, and drive alignment across senior stakeholders and technical teams. We are seeking a highly experienced Senior Technical Project Manager / Technical Program Manager to lead complex enterprise technology programs focused on cloud transformation, infrastructure modernization, IT operations, and AI-enabled operational initiatives .
Senior AI Platform & Agentic Infrastructure Engineer OKXSenior AI Platform & Agentic Infrastructure EngineerSan Jose, CA$178,000–$321,000 / yearWe assess demonstrated building in ways that respect your time and your confidentiality obligations to current and former employers: a portfolio deep-dive (walk us through systems you built and kept running, at the level of detail your obligations allow; we want architecture, decisions, and trade-offs, never proprietary code, data, or documents), a short, time-capped build exercise on a synthetic problem unrelated to OKX's business (a small agentic workflow with an evaluation harness, used for assessment only and never put to use by OKX; the work remains yours) that you defend live, walking us through your design decisions and extending it on the spot, a systems-design session on taking a prototype estate to production, and references focused on whether you built and operated systems in production. Build the agentic runtime and harness: orchestration, multi-model routing that sends each task to the right model, tool, MCP, and plugin integration across model providers, and the evaluation, red-team, and regression harnesses that grade agents before and after production, with evaluation gates, judge calibration, and cost controls that hold for each model, including prompt-injection and data-exfiltration threat modeling for agents that read untrusted content.
NewMember of Technical Staff, Cloud Infrastructure Fireworks.ai IncMember of Technical Staff, Cloud InfrastructureSan Mateo, NYFounded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed, NVIDIA, Sequoia Capital, and TCV, Fireworks powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. You'll partner closely with engineering partners, product teams, and infrastructure stakeholders to design solutions that balance performance, cost-efficiency, and operational simplicity across compute, storage, and networking layers.
NewSenior Product Marketing Manager, AI Training Infrastructure CoreWeave IncSenior Product Marketing Manager, AI Training InfrastructureSunnyvale, CA$177,000–$237,000 / yearThat means owning the competitive positioning framework field teams use against alternative infrastructure approaches, the persona-specific messaging that translates infrastructure capabilities into outcomes research leaders and platform teams care about, and the launch narrative for new capabilities as they reach the market. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. § 1157, or (iv) asylee under 8 U.S.C. § 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency.
Senior Technical Program Manager II, Infrastructure, Google Cloud GoogleSenior Technical Program Manager II, Infrastructure, Google CloudSunnyvale, CAFrom software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more. We're the driving force behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future.
Technical Program Manager III, Network Infrastructure, Google Cloud GoogleTechnical Program Manager III, Network Infrastructure, Google CloudSunnyvale, CAOur mission is to make it easy for Technical Infrastructure teams to develop and deliver products and execute the organizational goals by providing business operations and efficiency improvement solutions at scale to enhance organizational productivity, enable data driven business decisions. Utilize technical acumen, especially in networking hardware and software (switches/fabric), to drive project delivery, lead technology reviews, guide technical proposals, and build consensus across cross-functional teams.
Technical Program Manager, Infrastructure EtchedTechnical Program Manager, InfrastructureSan Jose, CaliforniaAs our Technical Program Manager, Infrastructure you will own end-to-end program delivery across the org — driving alignment between internal teams, managing external vendor relationships, and ensuring that every dependency, risk, and milestone is visible to the people who need to act on it. You'll be the connective tissue between engineering and leadership — translating technical complexity into clear status, surfacing risks before they become blockers, and facilitating resolution of the cross-disciplinary issues that emerge at the intersection of ASIC design, hardware platforms, and software integration.