Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract D-MatrixDirector, Site Reliability Engineering - AI Accelerator Infrastructure - ContractSanta Clara, CAOwn FinOps and capacity planning as a unified discipline across all three infrastructure tiers-cloud (AWS, Azure, GCP), colocation, and on-premises: establish spend visibility and attribution across every tier, model TCO comparatively, drive workload placement decisions based on cost and performance, and anticipate infrastructure needs for new silicon programs and customer deployments. Deep Linux systems expertise: networking (TCP/IP, RDMA, and bonding), kernel tuning, and bare-metal operations; hands-on experience with enterprise shared storage platforms (NAS/SAN, NFS/SMB at scale, and snapshot and replication architectures) and hybrid-cloud storage integration across on-prem and cloud tiers.
Staff AI Infrastructure Engineer BiohubStaff AI Infrastructure EngineerRedwood City, California$241,000–$331,000 / yearDesign and execute GPU cluster scaling plans, systematically validating storage, networking, interconnect, and scheduler behavior as clusters grow to support larger training runs. The technical problems are genuinely hard (e.g., multi-node distributed training, InfiniBand fabrics, large-scale storage, Slurm at scale) inside an organization where the work is aimed at helping people, not optimizing ad revenue.
AI Training Infrastructure Engineer – Humanoid Whole Body Control FigureAI Training Infrastructure Engineer – Humanoid Whole Body ControlSan Jose, CaliforniaThis role sits at the intersection of robotics, machine learning, controls, and software systems engineering, and is critical to how quickly we can iterate, train, and deploy new capability to our fleet of humanoid robots. Experience building or scaling training infrastructure for robotics, control systems, or large-scale ML workloads.
Software Engineer, Data Infrastructure Otter.aiSoftware Engineer, Data InfrastructureMountain View, CaliforniaUsing artificial intelligence, Otter generates real-time automated meeting notes, summaries, and other insights from in-person and virtual meetings - turning meetings into accessible, collaborative, and actionable data that can be shared across teams and organizations. If you're excited about building the backbone of a data-driven organization and thrive in a space where engineering excellence meets real-world impact, we'd love to talk.
Senior AI Infrastructure Engineer - Model Training KodiakSenior AI Infrastructure Engineer - Model TrainingMountain View, CA$190,000–$260,000 / yearShould the position require, and Kodiak determines that a candidate's residence, U.S. person status, and/or citizenship status necessitate an export license, bar the candidate from the position, or otherwise fall under national security-related restrictions, Kodiak will consider the candidate for alternative positions unaffected by such restrictions, under terms and conditions set forth at Kodiak's sole discretion, or, as an alternative, opt not to proceed with the candidate's application. Experience building high-performance data pipelines for large-scale training, including streaming dataset formats (WebDataset, MosaicML Streaming/MDS, or similar), sharding, and storage/network-aware loading.
Lightning Infrastructure Engineer (Remote) Lightning LabsLightning Infrastructure Engineer (Remote)Palo Alto, CaliforniaRemoteIn addition to the core Lightning Network Daemon software and end-user applications, the Lightning Network ecosystem includes supporting systems like watchtowers (a form of backup), peer availability/network monitoring, advanced liquidity provisioning tools, automated channel management, and other services. These tools will lower the barrier to entry for operating routing nodes[1] and enable existing routing node operators to more effectively manage their infrastructure.
AI Training Infrastructure Engineer - Humanoid Whole Body Control FigureAI Training Infrastructure Engineer - Humanoid Whole Body ControlSan Jose, CA$150,000–$300,000 / yearThis role sits at the intersection of robotics, machine learning, controls, and software systems engineering, and is critical to how quickly we can iterate, train, and deploy new capability to our fleet of humanoid robots. Experience building or scaling training infrastructure for robotics, control systems, or large-scale ML workloads.
Senior Front-End Infrastructure Engineer ProdaptSenior Front-End Infrastructure EngineerSan Jose, CaliforniaRegularly liaise with Synopsys vendors and other EDA partners to coordinate tool support, resolve escalated issues, and stay updated on new developments, maintaining an open line of communication to ensure high-quality support for verification processes. Candidates should possess a specialized skill set that encompasses regression management, tool deployment, and extensive interaction with industry-leading EDA vendors.
IT Infrastructure Administrator Frank Rimerman and Co LLPIT Infrastructure AdministratorSan Jose, California$115,000–$130,000 / yearFull timeWorking closely with senior infrastructure staff, this role is responsible for maintaining a secure, reliable, and scalable IT environment while troubleshooting technical issues, supporting infrastructure projects, and helping improve operational efficiency. Microsoft, VMware, Cisco, Azure, or other relevant industry certifications are preferred (e.g., Microsoft Certified: Azure Administrator Associate, VMware VCP, Cisco CCNA).
NewDomain Tech Lead - Infrastructure Patching Automation (Mainframe & As/400) DeloitteDomain Tech Lead - Infrastructure Patching Automation (Mainframe & As/400)San Jose, CA$134,000–$265,000 / yearThis role requires deep expertise in mainframe and AS/400 (IBM i) environments to lead and deliver the patch management program across complex legacy and hybrid infrastructure estates, develop automation solutions that reduce manual effort and risk, and serve as a trusted technical advisor to enterprise clients - working closely with client stakeholders, project managers, and cross-functional delivery teams to plan, schedule, and execute patch cycles while ensuring compliance, stability, and security across mission-critical platforms. As a Domain Tech Lead on the AI & Engineering Hybrid Cloud Infrastructure team, you will own end-to-end technical delivery within a specific infrastructure domain on a large-scale infrastructure engineering engagement, with a focus on automating the patching lifecycle of multiple infrastructure areas.
NewStaff Engineer, Infrastructure Platforms Pacific BiosciencesStaff Engineer, Infrastructure PlatformsMenlo Park, CaliforniaThis role partners closely with Research, Bioinformatics, Software Engineering, Information Security, and IT Operations to design, build, automate, secure, and continuously improve infrastructure platforms spanning enterprise Linux, hybrid cloud, virtualization, enterprise storage, High Performance Computing (HPC), and modern platform services. The Staff Engineer, Infrastructure Platforms provides technical leadership for the engineering, automation, security, and operational excellence of PacBio's hybrid infrastructure platform supporting enterprise, research, and scientific computing workloads.
Alibaba Cloud-RDMA Ops Engineer - Computing Infrastructure Networking-Sunnyvale Alibaba Group Holding LtdAlibaba Cloud-RDMA Ops Engineer - Computing Infrastructure Networking-SunnyvaleSunnyvale, CA$104,400–$171,000 / yearThis role focuses on building and operatiing ultra-low latency, high-throughput networks using RDMA technologies to power next-generation computing workloads. If hired, employee will be in an "at-will position" and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.
Alibaba Cloud-Cloud Infrastructure - Site Reliability Engineer (SRE)-Sunnyvale Alibaba Group Holding LtdAlibaba Cloud-Cloud Infrastructure - Site Reliability Engineer (SRE)-SunnyvaleSunnyvale, CA$104,400–$171,000 / yearStrong scripting skills in Shell/Python and experience with Infrastructure as Code (IaC) tools (Terraform preferred).Minimum Qualification: Experience: Over 2 years of experience in distributed systems reliability engineering, familiar with high-availability architecture design, and proficient in at least one of Python, Go, or Java. Preferred Qualification: SRE Practices: Familiar with core SRE practices (incident review, error budgeting, chaos engineering) and experienced in building automated risk control systems.
Senior Systems Engineer - AI Infrastructure Clockwork.ioSenior Systems Engineer - AI InfrastructurePalo Alto, CA$150,000–$230,000 / yearClockwork is pioneering a software-driven approach to AI fabrics by delivering cross-stack observability to catch and quickly resolve problems, workload fault tolerance to keep jobs running through failures, and performance acceleration that dynamically routes and paces traffic to avoid congestion. In addition to cash compensation, this role is eligible to participate in the company's equity program , which may include stock options granted in accordance with the company's equity plan and subject to approval and applicable vesting schedules.
Senior Software Engineer, Core Infrastructure OracleSenior Software Engineer, Core InfrastructureSanta Clara, CAOracle Cloud Infrastructure Workflow is a Tier 0 service that is critical to the smooth functioning of ALL basic OCI services by enabling their execution of distributed, multi-step work in a fault tolerant manner. Delivers features and load/performance tests; leverages data plane platforms and distributed state tools for high-volume retrieval, storage, and processing; and reviews peers' implementations for scalability compliance.
Machine Learning Infrastructure Engineer Institute of Foundation ModelsMachine Learning Infrastructure EngineerSunnyvale, CaliforniaYou’ll work side-by-side with world-class researchers and engineers to: • Extend distributed training frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod) • Implement distributed optimizers from mathematical specs • Build robust config + launch systems across multi-node, multi-GPU clusters • Own experiment tracking, metrics logging, and job monitoring for external visibility • Improve training system reliability, maintainability, and performance • While much of the work will support large-scale pre-training, pre-training experience is not required. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.
NewSenior Software Engineer, Robot Data Infrastructure Galactic Resource Advancement MechanismSenior Software Engineer, Robot Data InfrastructurePalo Alto, CaliforniaDemonstrated ownership of a production pipeline incident that caused silent data loss, corrupt data, or an unavailable downstream dataset, including detection, root cause, recovery or backfill, and a test or monitor that prevented recurrence. Success means a model behavior can be traced through its dataset, run, software, calibration, commands, interventions, outcomes, and hardware state—and that dataset revisions remain reproducible rather than becoming ungoverned data volume.
NewSenior Research Engineer, Foundation Model Training Infrastructure NvidiaSenior Research Engineer, Foundation Model Training InfrastructureSanta Clara, CAWays to stand out from the crowd: Master's or PhD's degree in Computer Science, Robotics, Engineering, or a related field; Demonstrated Tech Lead experience, coordinating a team of engineers and driving projects from conception to deployment; Strong experience at building large-scale LLM and multimodal LLM training infrastructure; Contributions to popular open-source AI frameworks or research publications in top-tier AI conferences, such as NeurIPS, ICRA, ICLR, CoRL. What we need to see: Bachelor's degree in Computer Science, Robotics, Engineering, or a related field; 10+ years of full-time industry experience in large-scale MLOps and AI infrastructure; Proven experience designing and optimizing distributed training systems with frameworks like PyTorch, JAX, or TensorFlow.
Site Reliability Engineer - Hardware Infrastructure NvidiaSite Reliability Engineer - Hardware InfrastructureSanta Clara, CAAssist teams in responding to high severity incidents, driving root cause analysis, crafting high-quality postmortems, and developing post-incident corrective actions. At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability.
Senior Software Engineer, Agentic Robotics Infrastructure NvidiaSenior Software Engineer, Agentic Robotics InfrastructureSanta Clara, CAOwn CI/CD across the applications repos - spanning CMake, Bazel, and Python builds - including runners, Docker images, dataset provisioning, nightly builds, and release and docs pipelines. Create agent skills and automated workflows that teams adopt in their own repos, and support embedded agents on the robot (code-as-policy and related approaches) where latency, safety, and determinism constraints apply.