Member of Technical Staff - GPU Infrastructure Hyperbolic LabsMember of Technical Staff - GPU InfrastructureSan Francisco, CaliforniaYou'll work at the cutting edge of cloud infrastructure, building the core orchestration layer that enables our platform to deliver up to 75% cost savings compared to traditional cloud providers. Experience with storage and data infrastructure for AI/ML workloads, including object storage, high-IOPS block storage, and distributed file systems for training data and checkpoints.
Staff Slurm Cluster & HPC Engineer BitdeerStaff Slurm Cluster & HPC EngineerSan Jose, CACluster health and reliability engineering- Build the passive and active health-check system expected of a top-tier GPU cloud: prolog/epilog checks, LBNL NHC or equivalent, DCGM diagnostics, and detection of XID/SXID errors, ECC faults, PCIe errors, GPUs falling off the bus, IB/RoCE link flaps, and NCCL stalls - with automatic drain and job requeue. Slurm cluster architecture and lifecycle- Design, deploy, and operate production Slurm clusters on bare metal and VMs: slurmctld/slurmdbd high availability, slurmrestd, configless slurmd, SACK/MUNGE and JWT authentication, and rolling version upgrades on live clusters without losing running jobs.
Staff Research Engineer, Discovery Team AnthropicStaff Research Engineer, Discovery TeamSan Francisco, CAThis research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. Strong candidates should have familiarity with language model training, evaluation, and inference, be comfortable triaging research ideas and diagnosing problems and enjoy working collaboratively.
Senior Staff Software Engineer, DC Infrastructure Crusoe EnergySenior Staff Software Engineer, DC InfrastructureSan Francisco, CA$250,000–$300,000 / yearWe're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services. Own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success.
Staff Software Engineer, DC Infrastructure Crusoe EnergyStaff Software Engineer, DC InfrastructureSan Francisco, CA$215,000–$260,000 / yearWe're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services. Own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success.
Hrbp LambdaHrbpSan Francisco, CaliforniaOur investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove. Proactive Talent Strategy : Actively involved in all interview loops, specifically the culture interview, to ensure new hires are aligned with the team's values and bring the right skills to meet the business’s needs.
Senior ML Operations Engineer Early Warning Services, LLCSenior ML Operations EngineerSan Francisco, CA$118,000–$169,000 / yearOur Machine Learning Operations team enables our Data Scientists to be able to build and deploy innovative models while developing cutting edge, cloud native capabilities to deliver predictive modeling solutions faster, more accurate, and more efficiently to help keep fraud and bad actors out of the banking system. Early Warning Services takes into consideration a variety of factors when determining a competitive salary offer, including, but not limited to, the job scope, market rates and geographic location of a position, candidate's education, experience, training, and specialized skills or certification(s) in relation to the job requirements and compared with internal equity (peers).
Staff Systems Engineer (Core Platform) General MotorsStaff Systems Engineer (Core Platform)Mountain View, CA$160,200–$263,700 / yearAs a Staff Systems Engineer, you will be the bridge between the platform and the hardware it manages -- building integration tooling, hardware discovery services, provisioning workflows, and agent software in Go that connects diverse automotive test equipment to the platform's control plane. Work with tools and technologies including Go, PostgreSQL, Nomad, Consul, Linux system interfaces, automotive communication protocols (CAN, Ethernet, serial), CI/CD pipelines, observability frameworks (Prometheus, Grafana, Datadog), and Git/GitHub.
Senior Channel Enablement Manager FastlySenior Channel Enablement ManagerSan Francisco, CARemote$129,470–$155,364 / yearLMS & Tool Proficiency: Hands-on experience administering Partner Relationship Management (PRM) systems, Learning Management Systems (LMS), and virtual facilitation tools (e.g., Zoom, Salesforce, WorkRamp, Highspot). Technical Aptitude: Ability to translate complex cloud technologies (e.g., edge cloud computing, CDN, API security, and serverless architectures) into clear, value-driven sales messaging for non-technical channel reps.
NewSenior Solutions Architect, Cluster Design And Architecture - Networking NvidiaSenior Solutions Architect, Cluster Design And Architecture - NetworkingSanta Clara, CAIn this role, you will be at the forefront of assisting with designs and architectures for next-generation networking solutions that connect thousands of GPUs and enable the world's most advanced AI supercomputers and enterprise AI infrastructure in the field. What you'll be doing: Partner with internal engineering efforts in GPU cluster building and networking and convey architecture and guidelines information both direct to customer and with field teams supporting customers.
Staff Advanced Concepts Optimization Engineer ArcherStaff Advanced Concepts Optimization EngineerSan Jose, CA$130,000–$160,000 / yearOur team develops multi-disciplinary physics simulations, advanced optimization methods, and large-scale cloud computing infrastructure, then deploys these tools to design optimal vehicle configurations. Archer is an aerospace company based in San Jose, California building an all-electric vertical takeoff and landing aircraft with a mission to advance the benefits of sustainable air mobility.
Senior Director, Applied Research Capital OneSenior Director, Applied ResearchSan Francisco, CaliforniaBasic Qualifications: PhD in Electrical Engineering, Computer Engineering, Computer Science, AI, Mathematics, or related fields plus 6 years of experience in Applied Research or M.S. in Electrical Engineering, Computer Engineering, Computer Science, AI, Mathematics, or related fields plus 8 years of experience in Applied Research. Key Responsibilities: Partner with a cross-functional team of scientists, machine learning engineers, software engineers, and product managers to deliver AI-powered platforms and solutions that change how customers interact with their money.
Vice President, Information Technology Procept BioroboticsVice President, Information TechnologySan Jose, CA$280,590–$330,105 / yearIt continues with our one-of-a-kind management program designed to build the best managers in the industry, where our people managers across functions come together to exchange ideas and grow, as both managers and learners, in an environment that challenges, supports and broadens. It starts with our live induction program that serves as an incubator for cross-functional team building, an immersion in Procept's history, jam-packed interactive sessions with executive leadership and a crash-course in the mission and purpose of what we do.
Senior Backend Java Engineer eTeam Inc.Senior Backend Java EngineerSunnyvale, CAThe ideal candidate will have strong expertise in Java-based microservices architecture, real-time data processing, distributed databases, containerization technologies, and cloud-native application development. The candidate will play a critical role in building resilient backend platforms capable of processing large-scale data streams while ensuring reliability, scalability, and performance.
NewSenior Software Engineer, Fleet Intelligence Backend NVIDIA CorpSenior Software Engineer, Fleet Intelligence BackendSanta Clara, CAWhat You'll Be Doing: This role focuses on backend services for GPU health, Fleet Intelligence, telemetry ingestion, inventory, attestation, alerting, reporting, and cloud operations automation, in addition, you will: Design and develop Go backend services, REST APIs, and data models for GPU Health and Fleet Intelligence. We are looking for a Senior Software Engineer to join our DGX Cloud / Fleet Intelligence team and build backend systems that power GPU health monitoring, telemetry ingestion, operational automation, and fleet visibility for NVIDIA's high-performance GPU infrastructure.
Senior Director, NCP And ISV Business Development NvidiaSenior Director, NCP And ISV Business DevelopmentSanta Clara, CA18+ overall years of experience in business development, strategic partnerships, alliances, cloud platforms, enterprise software, AI infrastructure, or a related field with 10+ years of leadership experience managing large, geographically distributed sales teams. Join NVIDIA, a trailblazer in computer graphics, AI, and advanced computational technologies, and embrace a pivotal role within our DGX Cloud and DSX OS teams as Senior Director, NCP & ISV Business Development.
Senior Manager, NCP And ISV Business Development NvidiaSenior Manager, NCP And ISV Business DevelopmentSanta Clara, CAJoin NVIDIA, a leader in computer graphics, AI, and accelerated computing, and become a key contributor to our DGX Cloud and DSX OS teams as a Senior Manager, NCP & ISV Business Development. Ways to stand out from the crowd: Experience working with NCPs, cloud infrastructure providers, enterprise ISVs, AI-native companies, or developer-platform partners.
Senior Product Dev Rel Engineer - DSX NvidiaSenior Product Dev Rel Engineer - DSXSanta Clara, CAArchitecting and publishing high-impact technical assets, including production-grade deployment guides, containerized workflows, reference architectures, and end-to-end DSX benchmarks showcasing efficient performance across NVIDIA GPU clusters. What you'll be doing: Driving developer-centric product improvements for DSX by collecting and synthesizing technical feedback, system-level friction, and workflow requirements from AI practitioners, enterprise data platform teams, and open-source communities into DSX Product Management and Core Engineering roadmaps.
NewNCX Senior Engineer NVIDIA CorpNCX Senior EngineerSanta Clara, CAProven experience collaborating with NVIDIA Cloud Partners, hyperscale cloud providers, managed AI clouds, or extensive service-provider infrastructure and operating SLOs for large-scale compute infrastructure and using operational data to improve availability, performance, and fleet efficiency. Extensive knowledge of infrastructure observability tools including Prometheus, Grafana, OpenTelemetry, Alertmanager, and scalable telemetry pipelines and translating reference architectures or infrastructure requirements into repeatable production operating models across multiple customer or partner environments.
NewSr Staff Software Engineer Palo Alto Networks IncSr Staff Software EngineerCA$145,000–$235,500 / yearFor candidates who receive an offer at the posted level, the starting base salary (for non-sales roles) or base salary + commission target (for sales/com-missioned roles) is expected to be the annual range listed below. Shifting from passive financial monitoring to AI-driven, self-optimizing architectures, you will focus heavily on AI/ML compute infrastructure like GPUs and LLM pipelines.