MLOps & Agentic Platform Engineer (AI Infrastructure) Hyphen Connect LimitedMLOps & Agentic Platform Engineer (AI Infrastructure)San Francisco, CaliforniaThe ideal candidate will have a strong DevOps/MLOps background and be adept at deploying scalable microservices and building observability dashboards. This role involves managing model registries, developing continuous training loops, and implementing A/B testing infrastructure.
LLM Pre-training & Distributed Engineer (AI Infrastructure) Hyphen Connect LimitedLLM Pre-training & Distributed Engineer (AI Infrastructure)San Francisco, CaliforniaThe ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure.
Infrastructure Engineering Advisor- Hybrid CignaInfrastructure Engineering Advisor- HybridWalnut Creek, CA$109,970–$176,300 / yearMust have experience with: performing Kafka cluster build, including design, infrastructure planning, High Availability and Disaster Recovery; Implementing wire encryption using SSL, authentication using SASL/LDAP and authorization using Kafka ACLs in Zookeeper; Creating Kafka connect cluster and connector deployment; Schemas and schema registry; Maintaining current Kafka clusters by performing upgrade, performance tuning, monitoring setup, and adding new features; Researching new features and conducting proof of concept installation and testing; Supporting Kafka in On-Premises and Cloud AWS; Kafka implementations and issue triage, including gathering requirements, understanding design, and assisting with build path to production, and performance and resiliency testing; remediating product vulnerabilities; Implementing best practices and working with product owners to implement roadmaps; Provisioning resources in AWS or Azure; Creating ec2 instances, EBS volumes, and security groups; Networking, load balancing, and okta integration; Automation using shell scripting, ansible, and terraform. eviCore Healthcare MSI, LLC seeks an Infrastructure Engineering Advisor for the Walnut Creek, CA location to design and develop Kafka-based data pipelines and messaging solutions to support real-time and near real-time data processing layer in the organization.
Member of Technical Staff, Infrastructure The Token CompanyMember of Technical Staff, InfrastructureSan Francisco, CaliforniaThe Token Company (YC W26, HF0 S26) is a seed stage startup in San Francisco, CA, and has raised ~$12M from First Round Capital, NEA, YC and SV Angel along with investors such as founders of Dropbox, Slack, Supercell and Huggingface and key people from OpenAI, xAI and DoorDash. You'd get to build global low-latency, high-throughput GPU ML inference infra that sits in the critical path of customer traffic, from deployment and scaling to reliability and cost-efficiency.
Autonomy Engineer - ML & DL Infrastructure Skydio, Inc.Autonomy Engineer - ML & DL InfrastructureSan Mateo, CA$170,000–$277,500 / yearThe Skydio team combines deep expertise in artificial intelligence, best-in-class hardware and software product development, operational excellence, and customer obsession to empower a broader, more diverse audience of drone users, from utility inspectors to first responders, soldiers in battlefield scenarios, and beyond. If you are excited about leveraging massive amounts of structured video data to solve open problems in object detection and tracking, optical flow estimation and segmentation, we would love to hear from you.
Research Scientist / Engineer – Reinforcement Learning Infrastructure LumaResearch Scientist / Engineer – Reinforcement Learning InfrastructureRedwood City, CaliforniaThe RL Infrastructure team builds the systems that make this possible at scale: high-throughput distributed training that couples policy optimization with large fleets of inference workers, environments that expose models to realistic multi-step tasks, and the reward, verification, and evaluation systems that turn model behavior into learning signal. Advance RL training efficiency and stability: sequence packing for long multi-turn trajectories, KV cache reuse across rollouts, curriculum and task sampling, and resource scheduling across heterogeneous training/inference workloads.
Principal Machine Learning Infrastructure Engineer, Ads & Discovery RobloxPrincipal Machine Learning Infrastructure Engineer, Ads & DiscoverySan Mateo, CA$295,250–$345,040 / yearYou are comfortable working across modern ML systems technologies such as FSDP, vLLM, SGLang, CUDA, distributed training frameworks, inference engines, and GPU kernels, while remaining tool-agnostic and focused on achieving step-function improvements in model quality, throughput, latency, reliability, and cost. Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators.
Senior Hardware Engineer - Infrastructure RobloxSenior Hardware Engineer - InfrastructureSan Mateo, CA$243,290–$295,250 / yearTechnical Knowledge: Strong understanding of Linux-based server environments, modern data center technologies, and server architecture including PCIe, NVMe, Ethernet/InfiniB and memory subsystems, and accelerator-based platforms. Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators.
Software Engineer: Infrastructure ThatchSoftware Engineer: InfrastructureSan Francisco, CaliforniaWe’re a fully distributed early stage company using technology to change the way America does healthcare. We hire engineers across multiple levels and care most about the scope you’ve owned, the impact you’ve had, and how you make decisions.
Principal Software Engineer, Data Infrastructure RobloxPrincipal Software Engineer, Data InfrastructureSan Mateo, CA$295,250–$345,040 / yearCross-Organizational Technical Leadership: A proven track record of influencing technical direction across a large engineering organization, leading consensus on complex architectural initiatives, and championing successful multi-quarter projects. You do not wait for clean specifications in high-ambiguity environments; you actively define the technical requirements, unblock dependencies, rally engineers across pods, and steer complex projects from initial whiteboard sketch all the way to production stability.
Forward Deployed Infrastructure Engineer Rethink recruitForward Deployed Infrastructure EngineerSan Francisco, New YorkThe Baseline team is a Forward Deployed Infrastructure Engineering team of roughly 90 people across the globe, ensuring seamless operations and fleet support across all environments — from on-prem to cloud, and from commercial to government networks. You will develop software and provide high-quality support for systems that are critical to solving the government's greatest challenges — across environments that range from commercial cloud to classified government networks.
Commercial Counsel-Infrastructure and GTM Together AICommercial Counsel-Infrastructure and GTMSan Francisco, CA$200,000–$230,000 / yearOur business runs on a large, multi-vendor compute footprint, and the terms we lock in early — covering capacity, pricing, service levels, and the freedom to change providers — shape our resilience for years. Lead negotiations for GPU and cloud capacity, colocation, power, data center space, networking, and related infrastructure-services agreements across our many suppliers.
Senior / Principal Infrastructure Engineer - ML Platform RobloxSenior / Principal Infrastructure Engineer - ML PlatformSan Mateo, CA$278,530–$345,040 / yearEvery day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators. Deep experience with Kubernetes (K8s) and cluster management at scale — e.g., managing 100s–1000s of nodes, serving 100k+ QPS, and ideally having experience writing custom Kubernetes controllers.
Member Of Technical Staff - Cloud Infrastructure SpaceXAIMember Of Technical Staff - Cloud InfrastructurePalo Alto, CA$180,000–$440,000 / yearWe are seeking a highly skilled Senior Infrastructure Engineer to join our US Government Team, focused on designing, building, and operating secure, scalable infrastructure for critical government projects. 5+ years of experience as an Infrastructure Engineer, Site Reliability Engineer, or similar role, with a focus on building and maintaining reliable, scalable systems, preferably in secure or government environments.
Associate, Infrastructure Strategy & Operations Together AIAssociate, Infrastructure Strategy & OperationsSan Francisco, CA$140,000–$170,000 / yearThe Associate, Infrastructure Strategy & Operations will be the analytical backbone of the Infrastructure Strategy team, powering the research, benchmarking, and operational analysis across the team's core workstreams, including capacity planning, compute sourcing, vendor evaluation, and site selection. Champion process improvements across the Infrastructure Strategy function, collaborating cross-functionally with Engineering, Data, and Finance to design AI-native workflows that streamline operations and automate repetitive analysis.
Hyperbolic Labs - Senior GPU Infrastructure Engineer deCircleHyperbolic Labs - Senior GPU Infrastructure EngineerSan Francisco, CaliforniaYou'll work at the cutting edge of cloud infrastructure, building the core orchestration layer that enables our platform to deliver up to 75% cost savings compared to traditional cloud providers. Experience with storage and data infrastructure for AI/ML workloads, including object storage, high-IOPS block storage, and distributed file systems for training data and checkpoints.
Business Development Lead - AI Infrastructure Hammerhead AIBusiness Development Lead - AI InfrastructureRedwood City, California7–10+ years of experience in technology and infrastructure focused business development, partnerships, or strategic sales, with a strong concentration on the infrastructure side of compute (cloud, data center, GPU/accelerator ecosystem). Our cutting-edge platform optimizes data center power infrastructure to maximize AI token generation within existing electrical limits, without requiring new power plants or grid expansions.
Staff Cloud Infrastructure Engineer AssuredStaff Cloud Infrastructure EngineerPalo Alto, CaliforniaThe challenges we face are deep and diverse, from creating digital experiences that provide comfort and clarity to claimants at their most stressed and vulnerable to orchestrating large-scale ML-driven decision-making on billions of dollars of claims payments, life at Assured is dynamic, collaborative, and rewarding. We’re looking for a Staff Cloud Infrastructure Engineer who’s passionate about automation, reliability, and empowering developers through world-class platform engineering.
Infrastructure Engineer TriumphInfrastructure EngineerSan Francisco, CaliforniaDrive our database strategy as a deep Postgres expert — query optimization, connection pooling, replication and read routing, schema evolution, and keeping our highest-traffic data paths healthy under load. You'll define how we build, ship, and scale from the ground up — CI/CD pipelines, the databases under our highest-traffic paths, and the telemetry that tells us what's actually happening in production.
ML Infrastructure Engineer SpaceXAIML Infrastructure EngineerPalo Alto, CA$180,000–$440,000 / year2+ years of industry experience working with high traffic or large-scale production environments, distributed systems, GPU infrastructure, and/or deep learning applications. RESPONSIBILITIES: Designing, building, and scaling GPU compute infrastructure, training frameworks, and experimentation tools to enable rapid iteration on ML hypotheses.