LLM Pre-training & Distributed Engineer (AI Infrastructure) Hyphen ConnectLLM Pre-training & Distributed Engineer (AI Infrastructure)San Francisco, CaliforniaThe ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure.
Senior Software Engineer, Core Infrastructure NooksSenior Software Engineer, Core InfrastructureSan Francisco, CaliforniaAround 40 engineers, mostly former founders or founding engineers and veteran builders from places like Applied Intuition, Scale, Airtable, Citadel, and Jane Street. Since then, we’ve tripled ARR each year for 3 years running and grown a high-caliber team turning the art of sales into a scalable science.
MLOps & Agentic Platform Engineer (AI Infrastructure) Hyphen ConnectMLOps & Agentic Platform Engineer (AI Infrastructure)San Francisco, CaliforniaThe ideal candidate will have a strong DevOps/MLOps background and be adept at deploying scalable microservices and building observability dashboards. This role involves managing model registries, developing continuous training loops, and implementing A/B testing infrastructure.
Principal Software Engineer, Money Infrastructure ReplitPrincipal Software Engineer, Money InfrastructureFoster City, CaliforniaWe’re looking for engineers who can design and scale reliable financial systems while translating complex monetary logic into intuitive, user-friendly experiences for both Replit customers and builders on the platform. Build high-converting, localized payment experiences across geographies — thinking globally while enabling users to pay locally.
Senior Backend Engineer (Infrastructure) SphereSenior Backend Engineer (Infrastructure)San Francisco, CaliforniaWe currently help companies like Lovable, Eleven Labs, Replit, Runway AI, Windsurf (+ many more) automate their sales tax, VAT and GST compliance: We're growing rapidly (30%+ MoM revenue growth) and helping the fastest growing tech companies around the world including Lovable, Eleven Labs, Replit, Windsurf, Deel, Linktree, Fal, Sierra AI etc. Work across our engineering organization to introduce and scale best practices with cloud-native technologies like Amazon ALB, ECS/EKS, Temporal, AWS SQS, Amazon Aurora PostgreSQL, Elasticache Redis, and S3.
Cloud Infrastructure Engineer SnowflakeCloud Infrastructure EngineerDublin, CaliforniaThese identity platforms run on cloud infrastructure (AWS and Azure) and Kubernetes, so comfort operating in a cloud environment is valuable — but the core mission of this role is identity: SSO, MFA, identity lifecycle management, access governance, and provisioning. Our team builds the authentication, authorization, identity governance, and identity lifecycle capabilities that keep Snowflake secure and compliant, partnering closely with Security and Engineering to deliver secure-by-default, Zero Trust access.
Software Engineer - Autonomy Infrastructure, Systems and Tools SkydioSoftware Engineer - Autonomy Infrastructure, Systems and ToolsSan Mateo, California$170,000–$240,000 / yearThe Skydio team combines deep expertise in artificial intelligence, best-in-class hardware and software product development, operational excellence, and customer obsession to empower a broader, more diverse audience of drone users, from utility inspectors to first responders , soldiers in battlefield scenarios , and beyond . Operate at the interface between autonomy software and core robotics middleware , defining clear APIs, data contracts, and performance requirements for messaging, state propagation, and coordination across subsystems, while partnering closely with downstream teams to support their implementation and integration.
Autonomy Engineer - Deep Learning Infrastructure SkydioAutonomy Engineer - Deep Learning InfrastructureSan Mateo, California$170,000–$236,500 / yearThe Skydio team combines deep expertise in artificial intelligence, best-in-class hardware and software product development, operational excellence, and customer obsession to empower a broader, more diverse audience of drone users, from utility inspectors to first responders , soldiers in battlefield scenarios , and beyond . If you are excited about leveraging massive amounts of structured video data to solve problems in Computer Vision (CV) such as object detection and tracking, optical flow estimation and segmentation, we would love to hear from you.
Autonomy Engineer - ML & DL Infrastructure SkydioAutonomy Engineer - ML & DL InfrastructureSan Mateo, California$170,000–$277,500 / yearThe Skydio team combines deep expertise in artificial intelligence, best-in-class hardware and software product development, operational excellence, and customer obsession to empower a broader, more diverse audience of drone users, from utility inspectors to first responders , soldiers in battlefield scenarios , and beyond . If you are excited about leveraging massive amounts of structured video data to solve open problems in object detection and tracking, optical flow estimation and segmentation, we would love to hear from you.
AI Agent Infrastructure Lead SphereAI Agent Infrastructure LeadSan Francisco, CaliforniaCreate workflows where agents can inspect the codebase, run local infrastructure, make changes, run tests, and prepare work for human review. Lead Sphere’s internal efforts to enable AI agents to act more autonomously across engineering, ops, customer success, tax research, and implementation.
ML Infra - Data Infrastructure The Bot CompanyML Infra - Data InfrastructureSan Francisco, CaliforniaStrong experience building and maintaining production systems that process PBs of data on thousands of machines. Throughout the process, we expect candidates to demonstrate: Exceptional mental acuity: you think quickly, learn instantly, and reason across unfamiliar domains.
Research Scientist / Engineer – Training Infrastructure LumaResearch Scientist / Engineer – Training InfrastructureRedwood City, CaliforniaThe Training Infrastructure team at Luma is responsible for building and maintaining the distributed systems that enable training of our large-scale multimodal models across thousands of GPUs. Research and implement advanced parallelization techniques (FSDP, Tensor Parallel, Pipeline Parallel, Expert Parallel).
Software Engineer - Autonomy Infrastructure, Systems And Tools Skydio, Inc.Software Engineer - Autonomy Infrastructure, Systems And ToolsSan Mateo, CA$170,000–$240,000 / yearThe Skydio team combines deep expertise in artificial intelligence, best-in-class hardware and software product development, operational excellence, and customer obsession to empower a broader, more diverse audience of drone users, from utility inspectors to first responders, soldiers in battlefield scenarios, and beyond. Operate at the interface between autonomy software and core robotics middleware, defining clear APIs, data contracts, and performance requirements for messaging, state propagation, and coordination across subsystems, while partnering closely with downstream teams to support their implementation and integration.
Autonomy Engineer - Deep Learning Infrastructure Skydio, Inc.Autonomy Engineer - Deep Learning InfrastructureSan Mateo, CA$170,000–$236,500 / yearThe Skydio team combines deep expertise in artificial intelligence, best-in-class hardware and software product development, operational excellence, and customer obsession to empower a broader, more diverse audience of drone users, from utility inspectors to first responders, soldiers in battlefield scenarios, and beyond. About the role: If you are excited about leveraging massive amounts of structured video data to solve problems in Computer Vision (CV) such as object detection and tracking, optical flow estimation and segmentation, we would love to hear from you.
MLOps & Agentic Platform Engineer (AI Infrastructure) Hyphen Connect LimitedMLOps & Agentic Platform Engineer (AI Infrastructure)San Francisco, CaliforniaThe ideal candidate will have a strong DevOps/MLOps background and be adept at deploying scalable microservices and building observability dashboards. This role involves managing model registries, developing continuous training loops, and implementing A/B testing infrastructure.
LLM Pre-training & Distributed Engineer (AI Infrastructure) Hyphen Connect LimitedLLM Pre-training & Distributed Engineer (AI Infrastructure)San Francisco, CaliforniaThe ideal candidate will have a deep understanding of GPU clusters and extensive experience in system engineering to ensure efficient and reliable training processes. This role is essential for orchestrating large-scale machine learning training runs and optimizing distributed infrastructure.
Infrastructure Engineering Advisor- Hybrid CignaInfrastructure Engineering Advisor- HybridWalnut Creek, CA$109,970–$176,300 / yearMust have experience with: performing Kafka cluster build, including design, infrastructure planning, High Availability and Disaster Recovery; Implementing wire encryption using SSL, authentication using SASL/LDAP and authorization using Kafka ACLs in Zookeeper; Creating Kafka connect cluster and connector deployment; Schemas and schema registry; Maintaining current Kafka clusters by performing upgrade, performance tuning, monitoring setup, and adding new features; Researching new features and conducting proof of concept installation and testing; Supporting Kafka in On-Premises and Cloud AWS; Kafka implementations and issue triage, including gathering requirements, understanding design, and assisting with build path to production, and performance and resiliency testing; remediating product vulnerabilities; Implementing best practices and working with product owners to implement roadmaps; Provisioning resources in AWS or Azure; Creating ec2 instances, EBS volumes, and security groups; Networking, load balancing, and okta integration; Automation using shell scripting, ansible, and terraform. eviCore Healthcare MSI, LLC seeks an Infrastructure Engineering Advisor for the Walnut Creek, CA location to design and develop Kafka-based data pipelines and messaging solutions to support real-time and near real-time data processing layer in the organization.
Member of Technical Staff, Infrastructure The Token CompanyMember of Technical Staff, InfrastructureSan Francisco, CaliforniaThe Token Company (YC W26, HF0 S26) is a seed stage startup in San Francisco, CA, and has raised ~$12M from First Round Capital, NEA, YC and SV Angel along with investors such as founders of Dropbox, Slack, Supercell and Huggingface and key people from OpenAI, xAI and DoorDash. You'd get to build global low-latency, high-throughput GPU ML inference infra that sits in the critical path of customer traffic, from deployment and scaling to reliability and cost-efficiency.
Autonomy Engineer - ML & DL Infrastructure Skydio, Inc.Autonomy Engineer - ML & DL InfrastructureSan Mateo, CA$170,000–$277,500 / yearThe Skydio team combines deep expertise in artificial intelligence, best-in-class hardware and software product development, operational excellence, and customer obsession to empower a broader, more diverse audience of drone users, from utility inspectors to first responders, soldiers in battlefield scenarios, and beyond. If you are excited about leveraging massive amounts of structured video data to solve open problems in object detection and tracking, optical flow estimation and segmentation, we would love to hear from you.
Research Scientist / Engineer – Reinforcement Learning Infrastructure LumaResearch Scientist / Engineer – Reinforcement Learning InfrastructureRedwood City, CaliforniaThe RL Infrastructure team builds the systems that make this possible at scale: high-throughput distributed training that couples policy optimization with large fleets of inference workers, environments that expose models to realistic multi-step tasks, and the reward, verification, and evaluation systems that turn model behavior into learning signal. Advance RL training efficiency and stability: sequence packing for long multi-turn trajectories, KV cache reuse across rollouts, curriculum and task sampling, and resource scheduling across heterogeneous training/inference workloads.