Principal Silicon Validation Engineer, Serdes/Pam4 Astera LabsPrincipal Silicon Validation Engineer, Serdes/Pam4San Jose, CA$185,000–$240,000 / yearIn this role you will formulate a comprehensive post-Silicon validation plan, automate the testing of ICs and board products, design experiments to root-cause unexpected behavior, report results and specification compliance, and work with key internal customers to quantify margins and ensure robustness. Astera Labs' Intelligent Connectivity Platform integrates CXL, Ethernet, NVLink, PCIe, and UALink semiconductor-based technologies with the company's COSMOS software suite to unify diverse components into cohesive, flexible systems that deliver end-to-end scale-up, and scale-out connectivity.
Senior Site Reliability Engineer Andromeda ClusterSenior Site Reliability EngineerSan Francisco, CaliforniaIncident Management: Proven track record leading incident response for complex distributed systems where the failure could be in hardware, firmware, networking, drivers, orchestration, or application code and you need to narrow it down fast. Observability & Monitoring: Hands-on experience building monitoring and alerting for GPU infrastructure, not just Prometheus/Grafana basics, but GPU-specific telemetry (DCGM, nvidia-smi, fabric manager metrics) integrated into actionable dashboards.
NewSenior Staff Software Engineer, SDN Architecture Crusoe EnergySenior Staff Software Engineer, SDN ArchitectureSan Francisco, CA$250,000–$300,000 / yearYou will design and build high-performance networking solutions optimized for AI and HPC workloads, working deeply across the Linux kernel, packet processing pipelines, and network virtualization to help deliver the extreme throughput and low latency that massive GPU clusters demand. We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.
Software Engineer - AI Infrastructure Andromeda ClusterSoftware Engineer - AI InfrastructureSan Francisco, CaliforniaPositioned at the intersection of infrastructure and product engineering, this role is deeply technical and systems-oriented, yet laser-focused on building solutions with broad leverage. Our job is to enable all of that fragmented compute to flow through one platform, delivering reliable capacity to model builders, research labs, and inference providers when they need it.
NewStaff Slurm Cluster & HPC Engineer BitdeerStaff Slurm Cluster & HPC EngineerSan Jose, CACluster health and reliability engineering- Build the passive and active health-check system expected of a top-tier GPU cloud: prolog/epilog checks, LBNL NHC or equivalent, DCGM diagnostics, and detection of XID/SXID errors, ECC faults, PCIe errors, GPUs falling off the bus, IB/RoCE link flaps, and NCCL stalls - with automatic drain and job requeue. Slurm cluster architecture and lifecycle- Design, deploy, and operate production Slurm clusters on bare metal and VMs: slurmctld/slurmdbd high availability, slurmrestd, configless slurmd, SACK/MUNGE and JWT authentication, and rolling version upgrades on live clusters without losing running jobs.
NewStaff AI Scheduling & Orchestration Engineer BitdeerStaff AI Scheduling & Orchestration EngineerSan Jose, CABitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. You will work at the intersection of distributed systems and AI, driving the architectural decisions that enable our platform to handle massive-scale distributed training and inference jobs with industry-leading efficiency.
Security Engineer, GRC PlaidSecurity Engineer, GRCSan Francisco, CaliforniaBuild Continuous Controls Monitoring: Automate evidence collection, control testing, and monitoring across cloud and internal systems, and write and tune the detection that flags drift and misconfiguration against baseline — so audit readiness is continuous and gaps surface the moment they appear, not at audit time. Compliance & risk knowledge: Working knowledge of SOC 2, ISO 27001/27701, and NIST CSF/800-53, with the ability to map controls to evidence and crosswalk a single control across frameworks.
NewSr. Cloud Support Engineer - Weekend Shift Crusoe EnergySr. Cloud Support Engineer - Weekend ShiftSan Francisco, CA$145,000–$175,000 / yearWe're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services. As a Cloud Support Engineer, you'll play a crucial role in empowering our customers to leverage this technology for groundbreaking advancements in fields like AI/ML, physics simulations, and computational biology.
NewSenior GPU Systems & Fabric Engineer BitdeerSenior GPU Systems & Fabric EngineerSan Jose, CAThis role requires deep expertise in Linux kernel internals, GPU architectures, and high-speed interconnects, as you will be tasked with transforming raw, bare-metal compute resources into scalable, resilient, and multi-tenant cloud primitives. Build and manage automated hardware remediation pipelines using DCGM telemetry to proactively identify, isolate, and reset degraded GPU/NIC components before they impact production jobs.
Staff Hardware Engineer- GNSS Connectivity General Motors CoStaff Hardware Engineer- GNSS ConnectivityMountain View, CA$160,000–$245,000 / yearWe are seeking a Staff Hardware Engineer GNSS to lead the definition, design, and validation of GNSS (Global Navigation Satellite System) solutions for our current and future generations of connectivity and compute platforms across a broad range of vehicles. This role will serve as the subject matter expert for GNSS hardware and system performance, working closely with product, systems, RF, hardware design, suppliers, and validation teams to deliver robust, automotive-grade GNSS solutions.
Staff Cluster Infrastructure Engineer ATOMSStaff Cluster Infrastructure EngineerSan Francisco, CaliforniaOur systems are designed to understand, predict, and control the real world with precision, turning complex physical operations into something more reliable, more scalable, and more productive. Our work only matters if it serves others, and we know that meaningful progress depends on the trust of the people we serve and the strength of our team — so we invest in both, creating an environment where you can do your best work and grow.
NewSenior Staff Software Engineer, Founding BMC Crusoe EnergySenior Staff Software Engineer, Founding BMCSunnyvale, CA$237,000–$288,000 / yearLead BMC bring-up on new server platforms - kernel, U-Boot, device tree, sensor management, fan and thermal control, power sequencing, host interfaces - working with partner engineering teams from schematics and hardware design docs. We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.
Senior Electrical Engineer ATOMSSenior Electrical EngineerSan Francisco, CaliforniaYou'll be working on cutting-edge technology that's transforming urban mobility and passenger transportation, designing robust hardware solutions for next-generation on-road autonomous vehicles. Our systems are designed to understand, predict, and control the real world with precision, turning complex physical operations into something more reliable, more scalable, and more productive.
NewSenior Full Stack Engineer ATOMSSenior Full Stack EngineerSan Francisco, California$176,000–$242,000 / yearOur systems are designed to understand, predict, and control the real world with precision, turning complex physical operations into something more reliable, more scalable, and more productive. We’re seeking a Full Stack Engineer to join our founding team who will design and build critical software interfaces that connect our physical AI to the digital world - keeping our applications reliable, secure and fast as we grow.
Senior New Production Introduction Engineer ATOMSSenior New Production Introduction EngineerSan Francisco, CaliforniaOur systems are designed to understand, predict, and control the real world with precision, turning complex physical operations into something more reliable, more scalable, and more productive. Our work only matters if it serves others, and we know that meaningful progress depends on the trust of the people we serve and the strength of our team—so we invest in both, creating an environment where you can do your best work and grow.
NewSenior AI/ML Engineer Condor SoftwareSenior AI/ML EngineerSan Francisco, CaliforniaYou will work close to real customer use cases as a core member of a cross-functional product team, collaborating with product managers, designers, quality engineers, and platform engineers to take AI capabilities from concept through production. Build, train, and serve custom ML models in production where they add value over general-purpose LLMs, owning the full lifecycle from data and training through deployment, monitoring, and retraining.
Security Engineer, Corporate Security NotionSecurity Engineer, Corporate SecuritySan Francisco, California$200,000–$220,000 / yearSecure AI tool usage at the endpoint, including governance of large language models, AI agents, and model context protocol (MCP) integrations; detect and prevent unauthorized or risky AI service access and data exfiltration through AI-enabled tools. This is a security engineering role focused on building scalable controls and automation across identity, endpoints, SaaS, and workforce infrastructure, not a traditional IT support or corporate engineering role.
Principal Engineer, Agentic AI Systems Inflection AIPrincipal Engineer, Agentic AI SystemsPalo Alto, CA$400,000–$550,000 / yearYou'll build production AI agents that can reason, retrieve context, use tools, and safely execute complex workflows while driving the design of scalable agent architectures, evaluation frameworks, observability, and reliability systems. 10+ years of engineering experience, including 3+ years at staff/principal or equivalent scope, and 2+ years building production-grade agentic or LLM-powered systems used by real customers.
Software Engineer, Internal Applications - Enterprise OpenAISoftware Engineer, Internal Applications - EnterpriseSan Francisco, CaliforniaRemoteFor unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. You will help reduce reliance on broadly privileged human actions, turn recurring technology problems into paved paths, and build agentic systems that can help resolve tickets end to end.
Staff Security Engineer, Infrastructure falStaff Security Engineer, InfrastructureSan Francisco, CaliforniaWe’re looking for a Security Engineer, Infrastructure to secure the core systems that power fal.ai’s platform: GPU compute, multi-cloud environments, networking, and data pipelines. You’ll operate across the full stack, from cloud and Kubernetes to identity, networking, and secrets, designing and implementing security controls that scale with a high-performance AI platform.