Software Engineer, GPU Infrastructure - HPC OpenAISoftware Engineer, GPU Infrastructure - HPCSan Francisco, CaliforniaFor unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. Experience with low level details of hardware components, protocols, and associated Linux tooling (e.g., PCIe, Infiniband, networking, power management, kernel perf tuning).
Senior Systems Engineer - AI Infrastructure Clockwork.ioSenior Systems Engineer - AI InfrastructurePalo Alto, CA$150,000–$230,000 / yearClockwork is pioneering a software-driven approach to AI fabrics by delivering cross-stack observability to catch and quickly resolve problems, workload fault tolerance to keep jobs running through failures, and performance acceleration that dynamically routes and paces traffic to avoid congestion. In addition to cash compensation, this role is eligible to participate in the company's equity program , which may include stock options granted in accordance with the company's equity plan and subject to approval and applicable vesting schedules.
NewAlibaba Cloud-RDMA Ops Engineer - Computing Infrastructure Networking-Sunnyvale Alibaba Group Holding LtdAlibaba Cloud-RDMA Ops Engineer - Computing Infrastructure Networking-SunnyvaleSunnyvale, CA$104,400–$171,000 / yearThis role focuses on building and operatiing ultra-low latency, high-throughput networks using RDMA technologies to power next-generation computing workloads. If hired, employee will be in an "at-will position" and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.
NewAlibaba Cloud-Cloud Infrastructure - Site Reliability Engineer (SRE)-Sunnyvale Alibaba Group Holding LtdAlibaba Cloud-Cloud Infrastructure - Site Reliability Engineer (SRE)-SunnyvaleSunnyvale, CA$104,400–$171,000 / yearStrong scripting skills in Shell/Python and experience with Infrastructure as Code (IaC) tools (Terraform preferred).Minimum Qualification: Experience: Over 2 years of experience in distributed systems reliability engineering, familiar with high-availability architecture design, and proficient in at least one of Python, Go, or Java. Preferred Qualification: SRE Practices: Familiar with core SRE practices (incident review, error budgeting, chaos engineering) and experienced in building automated risk control systems.
NewSenior Research Engineer, Foundation Model Training Infrastructure NvidiaSenior Research Engineer, Foundation Model Training InfrastructureSanta Clara, CAWays to stand out from the crowd: Master's or PhD's degree in Computer Science, Robotics, Engineering, or a related field; Demonstrated Tech Lead experience, coordinating a team of engineers and driving projects from conception to deployment; Strong experience at building large-scale LLM and multimodal LLM training infrastructure; Contributions to popular open-source AI frameworks or research publications in top-tier AI conferences, such as NeurIPS, ICRA, ICLR, CoRL. What we need to see: Bachelor's degree in Computer Science, Robotics, Engineering, or a related field; 10+ years of full-time industry experience in large-scale MLOps and AI infrastructure; Proven experience designing and optimizing distributed training systems with frameworks like PyTorch, JAX, or TensorFlow.
NewSenior Infrastructure Automation Engineer, Compute Platform NvidiaSenior Infrastructure Automation Engineer, Compute PlatformSanta Clara, CAWorking alongside our LSF internals engineer to encode hard-won scheduler knowledge into templates and policy, so that expertise lives in the repository instead of in one person's head. What you'll be doing: Designing and owning the configuration schema for LSF cell deployment, so that a policy change is written once, reviewed, tested, and applied identically everywhere it belongs.
Staff Infrastructure Engineer IvoStaff Infrastructure EngineerSan Francisco, CaliforniaRun Kubernetes like it's your own startup within the startup — own multi-cluster, multi-region deployments across AWS/GCP/Azure, with failover and disaster recovery that actually works when it matters (not just on paper). Office Perks: Enjoy a vibrant Downtown San Francisco office with catered lunch five days a week, premium snacks and coffee, an in-building gym, and a dog-friendly environment.
Infrastructure Engineer IvoInfrastructure EngineerSan Francisco, CaliforniaRun Kubernetes like it's your own startup within the startup — own multi-cluster, multi-region deployments across AWS/GCP/Azure, with failover and disaster recovery that actually works when it matters (not just on paper). Office Perks: Enjoy a vibrant Downtown San Francisco office with catered lunch five days a week, premium snacks and coffee, an in-building gym, and a dog-friendly environment.
Senior Infrastructure Engineer IvoSenior Infrastructure EngineerSan Francisco, CaliforniaRun Kubernetes like it's your own startup within the startup — own multi-cluster, multi-region deployments across AWS/GCP/Azure, with failover and disaster recovery that actually works when it matters (not just on paper). Office Perks: Enjoy a vibrant Downtown San Francisco office with catered lunch five days a week, premium snacks and coffee, an in-building gym, and a dog-friendly environment.
Senior Firmware Engineer - Development, Verification And Infrastructure NvidiaSenior Firmware Engineer - Development, Verification And InfrastructureSanta Clara, CAAs a member of our NVLink Firmware Development and Verification team, you will be responsible for performing unit and integration-level firmware verification across both pre-silicon and post-silicon platforms. Come, join our NVLink design team and help build the real-time, cost-effective computing platform driving our success in this exciting and quickly growing field.
Member of Technical Staff: Data Infrastructure & Data Operations Walden RoboticsMember of Technical Staff: Data Infrastructure & Data OperationsSan Francisco, CaliforniaAs a senior IC, you'll own the architecture and reliability at the core of that platform—ingestion, storage, transformation, and the data operations that keep quality high at scale—working hand in hand with the teams who produce and depend on that data. Participation in E-Verify does not limit your right to work and verification will only be completed after you become an employee with Walden Robotics.
Senior Mechanical Engineer, Test Infrastructure Form EnergySenior Mechanical Engineer, Test InfrastructureBerkeley, CAYou will be invested in developing and deploying novel, state-of-the-art test equipment including the design of mechanical and fluid handling systems, and electronics test stands supporting Hardware In the Loop (HIL) testing, while collaborating with the electrical subteam and supporting our operations team in keeping our tests online. Activities including designing, drafting, and manufacturing components & assemblies; engineering analysis including FEA, heat transfer, fluid mechanics, material selection, and manufacturability; communicating with stakeholders.
AI/ML Infrastructure Engineer ZensorsAI/ML Infrastructure EngineerSan Francisco, CaliforniaBuilding Efficient Operators: Working across the entire ML framework/compiler stack (e.g., PyTorch, CUDA, TensorRT, and NVIDIA DeepStream) to write custom optimized ML operator libraries. Optimizing Core ML Pipelines: Identifying key bottlenecks in our current video analytics pipeline and performing in-depth analysis to ensure the best possible performance on current server and edge compute architectures.
Solutions Architect - Infrastructure DecagonSolutions Architect - InfrastructureSan Francisco, CaliforniaOur technology enables industry-defining enterprises like Avis Budget Group, Block’s Cash App and Square, Chime, Oura Health, and Hunter Douglas to deploy AI agents that power personalized, deeply satisfying interactions across voice, chat, email, SMS, and every other channel. You will act as a trusted advisor to enterprise prospects and customers, guiding them through complex, ambiguous deployment conversations—from containerization strategies to multi-cloud architectures—ensuring Decagon integrates seamlessly into their existing infrastructure while meeting security, compliance, and scalability requirements.
AI Training Infrastructure Engineer – Humanoid Whole Body Control FigureAI Training Infrastructure Engineer – Humanoid Whole Body ControlSan Jose, California$150,000–$300,000 / yearThis role sits at the intersection of robotics, machine learning, controls, and software systems engineering, and is critical to how quickly we can iterate, train, and deploy new capability to our fleet of humanoid robots. Experience building or scaling training infrastructure for robotics, control systems, or large-scale ML workloads.
CI/CD and Automation Infrastructure Engineer Tanisha SystemsCI/CD and Automation Infrastructure EngineerCupertino, CAFull timeRole Purpose: Build and maintain scalable CI, automation and dashboard infrastructure that enables reliable test execution across media, platform and device validation workflows. • Build and maintain XCTest/XCUITest frameworks, Jenkins pipelines and supporting test infrastructure.
AI Training Infrastructure Engineer - Humanoid Whole Body Control FigureAI Training Infrastructure Engineer - Humanoid Whole Body ControlSan Jose, CA$150,000–$300,000 / yearThis role sits at the intersection of robotics, machine learning, controls, and software systems engineering, and is critical to how quickly we can iterate, train, and deploy new capability to our fleet of humanoid robots. Experience building or scaling training infrastructure for robotics, control systems, or large-scale ML workloads.
Software Engineer, Infrastructure - Analytics Platform OpenAISoftware Engineer, Infrastructure - Analytics PlatformSan Francisco, CaliforniaFor unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. This person should be able to own a platform such as CacheHouse end to end: define its technical direction, design its data model and storage architecture, integrate it with several research dashboards and workflows, guide one or two engineers, and ensure the system works reliably for its users.
Business Development Manager - AI Infrastructure & Servers ASUSTeK ComputerBusiness Development Manager - AI Infrastructure & ServersFremont, CA$170,000–$210,000 / yearEssential Duties and Responsibilities: Drive Vertical Growth: Expand market footprint across Cloud and Enterprise sectors, Including Next -Waved CSPs, Financial Services, Manufacturing, Retail, Healthcare, Energy & Utilities, Manufacturing and Telecommunications, Media & Entertainment, etc. As part of ASUS's evolution into a total solution provider, ISG delivers industry-leading servers, workstations, and software-defined solutions tailored for data centers, edge computing, and AI factories.
NewSenior Technical Program Manager IconmaSenior Technical Program ManagerPalo Alto, CA$78.19–$103.41 / hourDrive governance of non-human identities — service accounts, service principals, workload identities, and API credentials — spanning discovery, ownership assignment and attestation, least-privilege scoping, rotation, and decommissioning, with full lifecycle (joiner/mover/leaver) coverage for human and non-human accounts. You'll coordinate across enterprise security, infrastructure, compliance, legal, endpoint/digital workplace, and enterprise application teams; align our identity roadmap with the product security teams' own identity projects; and manage external implementation partners — keeping many threads moving, visible, and correctly sequenced against each other.