NewStaff Software Engineer, Storage Crusoe EnergyStaff Software Engineer, StorageSan Francisco, CA$240,000–$310,000 / yearWe're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services. Deep Performance Engineering: Lead "tiger teams" to solve the most ambiguous and difficult bottlenecks in the stack-from kernel-level IO context switching to global tail-latency in distributed clusters.
Principal PHY Serdes Validation Engineer MarvellPrincipal PHY Serdes Validation EngineerSanta Clara, CAThe Marvell post silicon validation group designs and develops test platforms for validating multi-core Arm-based Network processors and custom ASIC's, used in many communication infrastructure applications such as 5G base stations, NICs, Data Center and Cloud Computing platforms. Bachelor's degree in Computer Science, Electrical Engineering, or a related field and 10+ years of related professional experience, OR a Master's degree and/or PhD in Computer Science, Electrical Engineering, or a related field with 3-5 years of experience.
Manager, Technical Recruiting Crusoe EnergyManager, Technical RecruitingSan Francisco, CA$150,000–$200,000 / yearWe're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services. Reporting into the Head of Technical Recruiting, you'll be responsible for building, developing, and scaling a world-class technical recruiting team that hires exceptional talent across Engineering, Infrastructure, AI, Hardware, Product and other critical technical organizations.
NewCustomer Quality Engineer Arista NetworksCustomer Quality EngineerSanta Clara, CA$114,000–$167,000 / yearWe leverage the latest advancements in cloud computing, artificial intelligence, and software-defined networking to provide our clients with a competitive edge in an increasingly interconnected world. This role offers opportunities to collaborate with Engineering, New Product Engineering, Failure Analysis/Repair, Contract Manufacturers, Arista Manufacturing, Account teams, and Arista customers.
Senior/Staff Software Engineer - Infrastructure And Devops (Bay Area) FortanixSenior/Staff Software Engineer - Infrastructure And Devops (Bay Area)Santa Clara, CA$155,000–$230,000 / yearOur unified data security platform addresses vulnerabilities in hybrid multicloud environments, defends against threats, and makes it easier to discover, assess, and fix data exposure risks. Our commitment to solving the world’s toughest data security challenges has earned Fortanix multiple Cybersecurity Excellence and Innovation Awards, as well as recognition from industry giants such as NVIDIA, Microsoft, Intel, ServiceNow, and Snowflake.
Senior/Staff Software Engineer - Infrastructure and Devops (Bay Area) FortanixSenior/Staff Software Engineer - Infrastructure and Devops (Bay Area)Santa Clara, CAOur unified data security platform addresses vulnerabilities in hybrid multicloud environments, defends against threats, and makes it easier to discover, assess, and fix data exposure risks. Our commitment to solving the world’s toughest data security challenges has earned Fortanix multiple Cybersecurity Excellence and Innovation Awards, as well as recognition from industry giants such as NVIDIA, Microsoft, Intel, ServiceNow, and Snowflake.
NewInformation Technology (IT) Systems Administrator (Systems Application Analyst 3)-29952 Huntington Ingalls Industries IncInformation Technology (IT) Systems Administrator (Systems Application Analyst 3)-29952Mountain View, CA$96,278–$125,000 / yearThis role spans the full depth of IT administration: delivering expert-level Tier 1 and Tier 2 end-user support across multiple channels and platforms, managing enterprise network and cloud infrastructure, leading endpoint and mobile device management, and contributing directly to IT policy, procurement, and project planning. Serve as the primary administrator for cloud productivity platforms including Google Workspace and/or Microsoft 365, managing users, groups, licenses, permissions, and security configurations at an organizational level.
Sr. Customer Technical Architect LogicMonitorSr. Customer Technical ArchitectSan Francisco, CA$136,000–$190,000 / yearThe range displayed on each job posting reflects the minimum and maximum base salary target for new hires in the position, determined by work location and additional factors, including job-related skills, experience, interview performance, and relevant education or training. By anticipating issues before they strike, optimizing resources in real time, and enabling faster, smarter decisions, LogicMonitor helps IT and business leaders protect margins, accelerate innovation, and deliver exceptional digital experiences without compromise.
Senior GPU Infrastructure Engineer Hyperbolic LabsSenior GPU Infrastructure EngineerSan Francisco, CaliforniaYou'll work at the cutting edge of cloud infrastructure, building the core orchestration layer that enables our platform to deliver up to 75% cost savings compared to traditional cloud providers. Experience with storage and data infrastructure for AI/ML workloads, including object storage, high-IOPS block storage, and distributed file systems for training data and checkpoints.
Staff Software Engineer - Platform/Infrastructure Hyperbolic LabsStaff Software Engineer - Platform/InfrastructureSan Francisco, CaliforniaIn this role, you'll design the tenant management system, build resource lifecycle APIs, implement identity and access control, establish billing and quota frameworks, and create the multi-cloud abstractions that make our platform feel native regardless of the underlying provider. Deep experience building backend services in Go, including API servers, controllers/operators, resource lifecycle management systems, and distributed control plane components.
NewStaff Slurm Cluster & HPC Engineer BitdeerStaff Slurm Cluster & HPC EngineerSan Jose, CACluster health and reliability engineering- Build the passive and active health-check system expected of a top-tier GPU cloud: prolog/epilog checks, LBNL NHC or equivalent, DCGM diagnostics, and detection of XID/SXID errors, ECC faults, PCIe errors, GPUs falling off the bus, IB/RoCE link flaps, and NCCL stalls - with automatic drain and job requeue. Slurm cluster architecture and lifecycle- Design, deploy, and operate production Slurm clusters on bare metal and VMs: slurmctld/slurmdbd high availability, slurmrestd, configless slurmd, SACK/MUNGE and JWT authentication, and rolling version upgrades on live clusters without losing running jobs.
Senior Staff Software Engineer, DC Infrastructure Crusoe EnergySenior Staff Software Engineer, DC InfrastructureSan Francisco, CA$250,000–$300,000 / yearWe're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services. Own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success.
Staff Research Engineer, Discovery Team AnthropicStaff Research Engineer, Discovery TeamSan Francisco, CAThis research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. Strong candidates should have familiarity with language model training, evaluation, and inference, be comfortable triaging research ideas and diagnosing problems and enjoy working collaboratively.
NewStaff Software Engineer, DC Infrastructure Crusoe EnergyStaff Software Engineer, DC InfrastructureSan Francisco, CA$215,000–$260,000 / yearWe're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services. Own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success.
Tech Lead, AI Compute Infrastructure HeyGenTech Lead, AI Compute InfrastructurePalo Alto, CADevelop Large-Scale AI Job Framework: Build highly scalable, reliable frameworks for launching and managing massive, heterogeneous compute jobs, including multi-modal high-volume data ingestion/processing, distributed model training, and continuous evaluation/benchmarking. Optimize GPU Utilization: Design and implement mechanisms to aggressively optimize GPU and cluster utilization across thousands of devices for inference, training, data processing and large-scale deployment of our state-of-art video generation models .
Senior Cloud Engineer GridwareSenior Cloud EngineerSan Francisco, California$190,000–$210,000 / yearThe platform you’ll build and operate ingests millions of events per day from devices in the field, powers customer-facing dashboards and alerting, and supports the data science work that turns raw signals into grid intelligence. You will own AWS infrastructure, Kubernetes (EKS), CI/CD, and observability end-to-end, partnering with our Cloud Security team to keep the platform safe and compliant, and with backend, firmware, and data teams to keep them shipping fast.
Manager, Shared Services — Cloud Engineering - FedRAMP DelineaManager, Shared Services — Cloud Engineering - FedRAMPRedwood City, CaliforniaIn this role, you will own and evolve complex application infrastructure including messaging and data systems such as RabbitMQ, Elasticsearch, and Redis delivering reliable, scalable, and well-documented solutions that directly enable our development teams to move fast and build confidently. It is the only platform that enables you to discover all identities – including workforce, IT administrator, developers, and machines – assign appropriate access levels, detect irregularities, and respond to threats in real-time.
NewSolutions Architect, Infrastructure NvidiaSolutions Architect, InfrastructureSanta Clara, CAAs part of the NVIDIA Solutions Architecture team, we navigate uncharted technical and organizational spaces - serving as the bridge between early platform readiness, cloud engineering teams, product strategy, and large‑scale customer deployments. Proficiency with Linux systems tools for identifying issues and evaluating system performance, such as: dmesg, journalctl, lspci, numactl, ethtool, iostat, perf, nvidia-smi, top/htop, ipmitool, container‑level tooling, and related utilities.
Senior Manager, Cloud Engineering Governance - FedRAMP DelineaSenior Manager, Cloud Engineering Governance - FedRAMPRedwood City, CaliforniaServe as primary point of contact for cloud governance requests, escalations, and issues from Engineering and other departments; collect requirements and feedback when implementing new systems, guardrails, or CSP configurations; communicate policy changes and best practices to development teams. * Oversee multiple Microsoft Entra tenants used by Engineering and other departments, including cross-tenant synchronization, identity lifecycle management (provisioning, deprovisioning, attribute-based scoping), and SAML/OIDC authentication for SaaS applications and CSPs.
NewSenior/Staff Backend Software Engineer GridCARESenior/Staff Backend Software EngineerRedwood City, California$184,000–$284,000 / yearWhile leading tech companies invest billions in speculative, long-term solutions that may take decades to arrive, GridCARE’s pioneering physics-based generative AI platform unlocks gigawatts of hidden capacity in today’s electric grid — enabling hyperscalers, data center developers, and utilities to power AI infrastructure years sooner than conventional approaches and without costly upgrades. What matters most is depth: real experience in architecting, implementing, tuning, and debugging real time systems that have to run reliably in production, and genuine curiosity about growing that depth into broader platform and infrastructure ownership as GridCARE scales.