Senior Software Engineer, Cloud Infrastructure / SRE Oscar Health InsuranceSenior Software Engineer, Cloud Infrastructure / SRESan Francisco, CARemote$180,504–$236,911 / yearWe are responsible for architecting a world-class, resilient ecosystem using a modern stack centered on AWS/GCP, Terraform, and Kubernetes with developer focused tooling and CI/CD. Independently responsible for large or complex technology capabilities (set of components or services) within their team's domain or spanning multiple domains.
Senior Software Engineer - Machine Learning Infrastructure ATOMSSenior Software Engineer - Machine Learning InfrastructureSan Francisco, California$176,000–$230,000 / yearBuild and maintain software and tools for classical ML stack which helps data scientists manage full training lifecycle from data collection, data preparation, training, deployment and model serving. Our work only matters if it serves others, and we know that meaningful progress depends on the trust of the people we serve and the strength of our team—so we invest in both, creating an environment where you can do your best work and grow.
Senior Consultant, Built Environment and Infrastructure Control RisksSenior Consultant, Built Environment and InfrastructureSan Francisco, CAThe ideal candidate will have 5+ years of relevant experience and bring strong knowledge and understanding of CPTED principles and other global leading initiatives, including Secured by Design and recognized public and private sector safety and security planning guidance such as UFC, ISO31000, and NIST. The role includes undertaking fee-earning consultancy across a wide range of activities, including risk assessment, the definition of security strategies and high-level planning, concept and detailed master planning, and concept and detailed technical design.
Software Engineer - Training Infrastructure BaseTenSoftware Engineer - Training InfrastructureSan Francisco, CAFamiliarity or experience with the open source training stack and frameworks (NCCL, PyTorch, Megatron, NemoRL, VeRL, Axolotl, HF Trainer) and distributed training techniques (FSDP, DeepSpeed). By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production.
Technical Program Manager, Safeguards (Infrastructure & Evals) AnthropicTechnical Program Manager, Safeguards (Infrastructure & Evals)San Francisco, CAYour primary responsibility is driving reliability - owning the incident-response and post-mortem process, ensuring SLOs are defined and met in partnership with various teams, and making sure that when things go wrong, the right people know, the right actions get taken, and those actions actually get closed out. What You'll Do: Own the Safeguards Engineering ops review- Drive the recurring cadence that keeps the team informed and coordinated: surfacing recent incidents and failures, bringing visibility to reliability trends, and making sure the right people are in the room when decisions need to be made.
Software Engineer III, Embedded Systems/Firmware, AI and Infrastructure GoogleSoftware Engineer III, Embedded Systems/Firmware, AI and InfrastructureSunnyvale, CAFrom software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day.
Technical Program Manager, Infrastructure BaseTenTechnical Program Manager, InfrastructureSan Francisco, CAMaintain a clear-eyed view of cross-team dependencies and risks; surface them early, drive mitigation, and keep leadership informed with honest, signal-rich updates. The work is less about owning a single system and more about imposing order on ambiguity: standing up the right structures, driving decisions to closure, and making sure nothing falls through the cracks across dozens of stakeholders.
Member of Technical Staff - AI Cloud Infrastructure Emerald AIMember of Technical Staff - AI Cloud InfrastructureOakland, CaliforniaThey have built or served as a core early engineer on a managed cloud or AI platform, whether at a GPU cloud, an internal machine learning platform run at scale, a hyperscaler AI service, or a HPC research computing center operated as a service. Production experience deploying or operating Lustre or a comparable parallel filesystem such as GPFS, Weka, VAST, or BeeGFS, with a solid understanding of parallel filesystem architecture, tuning, and failure modes.
Founding Software Engineer - Backend and Infrastructure unsiloed.aiFounding Software Engineer - Backend and InfrastructureSan Francisco, CaliforniaWhat you Will Do As a Founding Software Engineer, you will own the entire technical stack end-to-end, from core infrastructure and production systems to deployment, reliability, and developer experience. Architect, build, and scale core backend systems powering document intelligence and VLM-based workflows.
Software Engineer III, Performance, AI and Infrastructure GoogleSoftware Engineer III, Performance, AI and InfrastructureSunnyvale, CAFrom software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day.
Applications Architect Manager UnisysApplications Architect ManagerSanta Clara, CA$120,000–$160,000 / yearBe knowledgeable about the assigned Service Tower(s), knowledgeable about other Service Provider Service Tower services that impact their assigned area(s), knowledgeable about Service Provider subcontractor and Third Party services and how all of these integrate to provide Services for the Customer. Knowledgeable on project management concepts including project scheduling, risk management, estimation (infrastructure estimation / IT Service Management estimation), Work Breakdown Structures, requirements management / traceability, Change Management, etc.
Manager, Infrastructure Capex Accounting AnthropicManager, Infrastructure Capex AccountingSan Francisco, CAIn this role, you will execute day-to-day accounting for Anthropic's owned infrastructure assets - the GPUs, networking equipment, and data center buildouts underpinning our recently announced $50bn investment in American computing infrastructure. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences.
Technical Program Manager, AI Infrastructure Character AITechnical Program Manager, AI InfrastructureRedwood City, CAThis role is a good match for someone who thrives in technically complex environments, enjoys bringing structure to ambiguous problem spaces, and can drive alignment across deeply cross-functional teams. In this role, you'll partner closely with engineering, research, and product teams to shape infrastructure strategy, align roadmaps, and drive end-to-end execution for critical initiatives across training, evaluation, and inference.
Director, Compute Infrastructure Procurement Operations AnthropicDirector, Compute Infrastructure Procurement OperationsSan Francisco, CAYou will simultaneously operationalize three enterprise-scale initiatives: a multi-entity procurement structure spanning domestic and international jurisdictions, a compliance-critical Asset Lifecycle Management program and procurement operations for Anthropic data center sites across the US with international expansion underway. Anthropic is seeking a Director of Compute Infrastructure Procurement Operations to build and lead the function responsible for how Anthropic procures, tracks, and governs its compute infrastructure assets at scale.
Infrastructure Product Manager ATOMSInfrastructure Product ManagerSan Francisco, California$224,000–$284,000 / yearDrive Product Strategy: Own the roadmap for our internal services and platforms, including Core Compute (Kubernetes), Networking (Service Mesh/Istio), extensive in-house storage primitives (Postgres, Redis/Valkey, Kafka, CRDB, etc.), DevEx/Observab/Costs platforms,, Data and ML platforms, Security Platform. In this role, you will act as a force multiplier for our entire engineering organization, overseeing a complex ecosystem that supports everything from core compute primitives to in-house built distributed systems such as KEQ , Cloudless Blob , Splitter , LogProc , and others.
Cloud Infrastructure Engineer Alchemy Insights, IncCloud Infrastructure EngineerSan Francisco, CADesign and manage multi-cloud, multi-region network architecture- VPC design, IPAM, DNS (Cloudflare), cross-cloud connectivity, security groups, and edge-proxy/istio gateway configuration. The Infrastructure team provides the infrastructure, tooling, and expertise needed to allow Alchemy engineers to ship, scale, and operate high-quality products in a fast, safe, and cost-efficient manner.
Staff Machine Learning Infrastructure Engineer ATOMSStaff Machine Learning Infrastructure EngineerSan Francisco, California$224,000–$280,000 / yearYou will own the challenge of scaling distributed GPU workloads to support a high volume of concurrent training runs across an expanding vehicle fleet, building a platform that can flexibly run on whatever GPU capacity is available, regardless of provider or environment, directly accelerating innovation across the platform. Distributed Computing & Orchestration: Leverage distributed compute frameworks to efficiently manage and execute a high volume of complex ML training jobs concurrently across large GPU clusters.
Senior Software Engineer, Infrastructure, Google Cloud Platforms GoogleSenior Software Engineer, Infrastructure, Google Cloud PlatformsSunnyvale, CAWe're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve.
NewEngineering Manager, GPU Infrastructure CohereEngineering Manager, GPU InfrastructureSan Francisco, CaliforniaA background running large Kubernetes compute fleets in production, including in multi-cloud environments: multi-cluster operations, scheduling, node health at scale, and familiarity with IaC and infrastructure monitoring. You’ve gone deep in one of the layers that make a GPU training fleet work, whether that’s cluster-wide operations, GPU networking, or hardware, and you’re willing to get hands-on and learn the rest.
Backend / Infrastructure Engineer Elite Talent ConsultingBackend / Infrastructure EngineerSan Francisco, CaliforniaIn this role, you will design and develop scalable infrastructure that powers our core product, enabling radiologists to work more efficiently and double their productivity. About the Role: We are hiring mid-level and senior backend engineers to help build an AI-powered platform that is revolutionizing workflows in radiology.