Machine Learning Infrastructure Engineer David Joseph & CompanyMachine Learning Infrastructure EngineerSan Francisco, California$200,000–$400,000 / yearTech stack: Distributed training frameworks (FSDP, DeepSpeed), NVIDIA GPUs, Linux, Python, C++, Kubernetes/Docker, and a major cloud platform (GCP, AWS, or Azure). Research and test training approaches, including parallelization techniques and numerical-precision trade-offs across model scales.
Director of Operations - IT Infrastructure & Server Rack Manufacturing AMAXDirector of Operations - IT Infrastructure & Server Rack ManufacturingFremont, CA$140,000–$180,000 / yearCatering to industries such as AI, cloud computing, autonomous vehicles, and high-performance computing, AMAX has set benchmarks in innovation, including pioneering liquid-cooled HPC systems for the semiconductor industry. Committed to addressing the growing demands of AI, AMAX delivers advanced solutions that help organizations achieve their technology goal and drive progress on a global scale.
Member of Technical Staff, Infrastructure LatchBioMember of Technical Staff, InfrastructureSan Francisco, CaliforniaYou will be working on problems including agent infrastructure and orchestration, ML infrastructure, model benchmarking at scale, full stack web development, developer and user experience, and process optimization. Infrastructure challenges span container orchestration, sandboxing, distributed filesystems, cloud architecture, and database systems.
NewInternship, Software Engineer, AI Data Infrastructure (Winter/Spring 2027) Tesla IncInternship, Software Engineer, AI Data Infrastructure (Winter/Spring 2027)Palo Alto, CAThe Telemetry team is responsible for the full lifecycle of this data: from specifying interesting events for data collection, to efficiently recording as much relevant data as possible on our embedded autopilot computer, to processing the data in the cloud. This allows us to gather data from our fleet of millions of vehicles around the world to feed the training of our Neural Networks, providing Tesla with a significant competitive advantage in the race to full autonomy.
Applied AI Health Data Architect-Senior Manager PwCApplied AI Health Data Architect-Senior ManagerSan Francisco, CA$124,000–$280,000 / yearExamples of the skills, knowledge, and experiences you need to lead and deliver value at this level include but are not limited to: Craft and convey clear, impactful and engaging messages that tell a holistic story. As a Senior Manager, you will influence key stakeholders, making sure that innovative data solutions meet the evolving needs of health systems while developing productive teams committed to excellence.
Product Manager, RF Portfolio - Embedded Technologies Advanced Micro Devices, IncProduct Manager, RF Portfolio - Embedded TechnologiesSan Jose, CaliforniaThe Product Line Manager will define and drive product strategy, market positioning, customer engagement, new product introduction readiness, and go-to-market execution for the Versal RF Series portfolio and Zynq UltraScale+ RFSoC product line. The successful candidate will thrive in a fast-paced environment, operate effectively across multiple cross-functional teams, and demonstrate the ability to translate customer requirements, market trends, and financial objectives into a cohesive and actionable product line strategy.
RF Portfolio Product Line Manager - Embedded Technologies Advanced Micro Devices IncRF Portfolio Product Line Manager - Embedded TechnologiesSan Jose, CAThe Product Line Manager will define and drive product strategy, market positioning, customer engagement, new product introduction readiness, and go-to-market execution for the Versal RF Series portfolio and Zynq UltraScale+ RFSoC product line. The successful candidate will thrive in a fast-paced environment, operate effectively across multiple cross-functional teams, and demonstrate the ability to translate customer requirements, market trends, and financial objectives into a cohesive and actionable product line strategy.
Infrastructure Engineer FlintInfrastructure EngineerSan Francisco, CaliforniaTechCrunch Imagine a website that builds new pages, tests creative, and optimizes performance in response to real-world signals: competitive shifts, user behavior, trending search queries, the way an autonomous vehicle responds to road conditions and traffic. - Dan Levine, Accel Flint is led by co-founders Michelle Lim , Engineer #1 turned Head of Growth and Product at Warp, and Max Levenson , engineering leadership at autonomous vehicle startup Nuro, then Engineer #2 at Vooma (YCS23).
Member of Technical Staff — Training Infrastructure Causal LabsMember of Technical Staff — Training InfrastructureSan Francisco, CaliforniaOur founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN. Research and test parallelization strategies and numerical precision trade-offs across model scales, including for architectures that don't map cleanly onto existing LLM training stacks.
Software Engineer, Data Infrastructure & Acquisition - Santa Clara, CA, USA SpeechifySoftware Engineer, Data Infrastructure & Acquisition - Santa Clara, CA, USASanta Clara, CAOver 50 million people use Speechify’s text-to-speech products to turn whatever they’re reading – PDFs, books, Google Docs, news articles, websites – into audio, so they can read faster, read more, and remember more. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and Google, leading PhD programs like Stanford, high growth startups like Stripe, Vercel, Bolt, and many founders of their own companies.
Member of Technical Staff, Infrastructure MirendilMember of Technical Staff, InfrastructureSan Francisco, CaliforniaTraining and inference infrastructure - understand the resource and scheduling demands of research, training, and inference workloads and build the platform capabilities those workloads need. Networking - build the networking layer across clouds, clusters, and hosts: routing, peering, load balancing, and network isolation.
Software Engineer, Data Infrastructure & Acquisition - Menlo Park, CA, USA SpeechifySoftware Engineer, Data Infrastructure & Acquisition - Menlo Park, CA, USAMenlo Park, CAOver 50 million people use Speechify’s text-to-speech products to turn whatever they’re reading – PDFs, books, Google Docs, news articles, websites – into audio, so they can read faster, read more, and remember more. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and Google, leading PhD programs like Stanford, high growth startups like Stripe, Vercel, Bolt, and many founders of their own companies.
Data Infrastructure Engineer Mind RoboticsData Infrastructure EngineerPalo Alto, CaliforniaAs a Data Infrastructure Engineer, you'll build and operate the pipelines that turn raw sensor and demonstration data into training-ready datasets — from ingestion off real robots and capture devices, through processing and quality filtering, to the dataloaders that feed model training. That starts with a data engine: the pipelines and infrastructure that turn raw, messy, multimodal sensor streams from robots and human demonstrations into high-quality, well-curated training data at scale.
Senior Infrastructure PM Hudson ManpowerSenior Infrastructure PMSan Jose, CaliforniaRemoteResponsibilities include solving complex technical problems, conducting feasibility studies, prioritizing project tasks, and collaborating closely with IT senior management. JOB DESCRIPTION: As a Project Management team member, the Infrastructure PM will focus on infrastructure and application initiatives, developing charters, plans, deliverables, and other project artifacts.
Software Engineer, ML Infrastructure Realm LabsSoftware Engineer, ML InfrastructureSunnyvale, CaliforniaYou will design and operate the full ML serving stack from model artifacts to GPU execution, and work closely with Product and ML teams to ensure our models can support high QPS, strict SLAs, and production correctness . You will apply deep knowledge of model internals to deploy, optimize, and run modern LLMs at scale , owning performance end-to-end across latency, throughput, and reliability .
HPC/ML Infrastructure Engineer SpellbrushHPC/ML Infrastructure EngineerSan Francisco, CaliforniaYou’ll serve as the bridge between our researchers and the bare GPU machines, helping to make sure that SLURM jobs are running, parallel filesystems are serving, network is transmitting, and that the anime models are training. Now the landscape has become unbelievable complex, with SLURM deploys through Slinky on K8s, provisioning through warewulf/MAAS/ansible, filesystems through WEKA/VAST/Ceph, VPN and access through tailscale, and monitoring via the Grafana/Prometheus stack.
Member of Technical Staff — Inference Infrastructure Causal LabsMember of Technical Staff — Inference InfrastructureSan Francisco, CaliforniaOur founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN. Progress on an LPM is gated by how fast we can evaluate it: large-scale backtesting against decades of physical observations, ensemble generation, and rollout evaluation across model scales.
AI Infrastructure Engineer SpellbrushAI Infrastructure EngineerSan Francisco, CaliforniaSpellbrush, the world’s leading generative AI studio behind nijijourney , is looking for an AI Infrastructure Engineer to join us in building out end-to-end ML infrastructure to run our models on all platforms. Work alongside a fast-paced and nimble team developing the latest state-of-the-art image generation models serving over 16 million users.
Member of Technical Staff — Data Infrastructure Causal LabsMember of Technical Staff — Data InfrastructureSan Francisco, CaliforniaOur founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN. Your mission is to build the data platform underneath it all — the storage, compute, and loading systems that make every dataset cheap to ingest, fast to query, and immediately available to training.
Software Engineer, ML Infrastructure, Level 4 SnapchatSoftware Engineer, ML Infrastructure, Level 4Palo Alto, CA$157,000–$235,000 / yearThe Company operates Snapchat, a visual messaging app that enhances your relationships with friends, family, and the world, and Specs Inc., a wholly-owned subsidiary dedicated to making computing more human, in addition to Bitmoji, Saturn, and other digital services. 2+ years of post-Bachelor's software development experience; or Master's degree in a technical field + 1+ year of post-grad software development experience; or PhD in a relevant technical field.