Mechanical Engineer, Test Infrastructure BelcanMechanical Engineer, Test InfrastructureBerkeley, caDerisking & Validation: Work closely with the Design Engineering and Test teams to translate high-level testing requirements into functional mechanical hardware that identifies failure modes early in the development cycle. * System-Level Integration: Create test environments that accurately simulate real-world conditions for our battery packs and auxiliary systems (pumps, manifolds, and power electronics).
Machine Learning Infrastructure Engineer Institute of Foundation ModelsMachine Learning Infrastructure EngineerSunnyvale, CaliforniaYou’ll work side-by-side with world-class researchers and engineers to: • Extend distributed training frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod) • Implement distributed optimizers from mathematical specs • Build robust config + launch systems across multi-node, multi-GPU clusters • Own experiment tracking, metrics logging, and job monitoring for external visibility • Improve training system reliability, maintainability, and performance • While much of the work will support large-scale pre-training, pre-training experience is not required. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.
Manager, Workplace Projects & Infrastructure VOLEONManager, Workplace Projects & InfrastructureBerkeley, CAReporting to the Managing Director, Strategic Projects, the ideal candidate is a highly organized, execution-oriented project manager with experience managing complex workplace initiatives and driving results across multiple stakeholders. Proven ability to independently manage complex projects involving multiple stakeholders, budgets, timelines, and vendors, and drive successful outcomes in fast-paced and evolving environments.
IT Facilities & Infrastructure Project Manager : 26-01934 Akraya Inc.IT Facilities & Infrastructure Project Manager : 26-01934Sunnyvale, CA$80–$85 / hourWe are seeking an IT Facilities & Infrastructure Project Manager to lead the delivery of enterprise IT infrastructure and facilities projects, including network infrastructure, structured cabling, AV systems, physical security, and data center initiatives. The ideal candidate will have extensive project management experience, strong stakeholder and vendor management skills, and the ability to successfully deliver complex, high-visibility infrastructure projects in a fast-paced environment.
Staff Infrastructure Engineer IvoStaff Infrastructure EngineerSan Francisco, CaliforniaRun Kubernetes like it's your own startup within the startup — own multi-cluster, multi-region deployments across AWS/GCP/Azure, with failover and disaster recovery that actually works when it matters (not just on paper). Office Perks: Enjoy a vibrant Downtown San Francisco office with catered lunch five days a week, premium snacks and coffee, an in-building gym, and a dog-friendly environment.
NewSenior Software Engineer, Robot Data Infrastructure Galactic Resource Advancement MechanismSenior Software Engineer, Robot Data InfrastructurePalo Alto, CaliforniaDemonstrated ownership of a production pipeline incident that caused silent data loss, corrupt data, or an unavailable downstream dataset, including detection, root cause, recovery or backfill, and a test or monitor that prevented recurrence. Success means a model behavior can be traced through its dataset, run, software, calibration, commands, interventions, outcomes, and hardware state—and that dataset revisions remain reproducible rather than becoming ungoverned data volume.
Program Manager, Operations Infrastructure Artech LLCProgram Manager, Operations InfrastructureFoster City, CA$53–$57.10 / hourWe need a highly capable, forward-thinking, and tenacious Program Manager to drive the deployment of our operations infrastructure projects (focused on IT and data infrastructure) as well as support autonomous navigation within these sites. Collaborate cross-functionally to bridge technical requirements with operational needs.
Senior Infrastructure Engineer - AI Open SelectSenior Infrastructure Engineer - AISan Francisco, CaliforniaYou’ll tackle challenges at the intersection of cloud infrastructure, telephony, and machine learning - building systems that handle high-throughput, low-latency workloads and integrating with enterprise-grade voice and video systems. As a Senior Infrastructure Engineer, you’ll architect and scale distributed systems that power millions of real-time, AI-driven phone conversations for major brands and enterprises.
Staff Software Engineer, Data Infrastructure PatreonStaff Software Engineer, Data InfrastructureSan Francisco, CaliforniaFamiliarity with cloud infrastructure and broad fluency across data stores: Object storage (S3), RDMS (MySQL, or Postgres, etc.), Key value (DynamoDB), OLAP (Clickhouse, Pinot, or Druid, etc.), search/indexing (Elasticsearch) plus data flow patterns for keeping them in sync with source systems. You'll join a small, high-craft team that partners across Product, Data Science, Infrastructure, and the rest of Engineering to make sure the platform underneath all of that work is reliable, scalable, and easy for other engineers to build on.
Member of Technical Staff: Data Infrastructure & Data Operations Walden RoboticsMember of Technical Staff: Data Infrastructure & Data OperationsSan Francisco, CaliforniaAs a senior IC, you'll own the architecture and reliability at the core of that platform—ingestion, storage, transformation, and the data operations that keep quality high at scale—working hand in hand with the teams who produce and depend on that data. Participation in E-Verify does not limit your right to work and verification will only be completed after you become an employee with Walden Robotics.
Technical Program Manager, Safeguards (Infrastructure & Evals) AnthropicTechnical Program Manager, Safeguards (Infrastructure & Evals)San Francisco, CAYour primary responsibility is driving reliability - owning the incident-response and post-mortem process, ensuring SLOs are defined and met in partnership with various teams, and making sure that when things go wrong, the right people know, the right actions get taken, and those actions actually get closed out. What You'll Do: Own the Safeguards Engineering ops review- Drive the recurring cadence that keeps the team informed and coordinated: surfacing recent incidents and failures, bringing visibility to reliability trends, and making sure the right people are in the room when decisions need to be made.
NewSenior Research Engineer, Foundation Model Training Infrastructure NvidiaSenior Research Engineer, Foundation Model Training InfrastructureSanta Clara, CAWays to stand out from the crowd: Master's or PhD's degree in Computer Science, Robotics, Engineering, or a related field; Demonstrated Tech Lead experience, coordinating a team of engineers and driving projects from conception to deployment; Strong experience at building large-scale LLM and multimodal LLM training infrastructure; Contributions to popular open-source AI frameworks or research publications in top-tier AI conferences, such as NeurIPS, ICRA, ICLR, CoRL. What we need to see: Bachelor's degree in Computer Science, Robotics, Engineering, or a related field; 10+ years of full-time industry experience in large-scale MLOps and AI infrastructure; Proven experience designing and optimizing distributed training systems with frameworks like PyTorch, JAX, or TensorFlow.
Site Reliability Engineer - Hardware Infrastructure NvidiaSite Reliability Engineer - Hardware InfrastructureSanta Clara, CAAssist teams in responding to high severity incidents, driving root cause analysis, crafting high-quality postmortems, and developing post-incident corrective actions. At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability.
Senior Firmware Engineer - Development, Verification And Infrastructure NvidiaSenior Firmware Engineer - Development, Verification And InfrastructureSanta Clara, CAAs a member of our NVLink Firmware Development and Verification team, you will be responsible for performing unit and integration-level firmware verification across both pre-silicon and post-silicon platforms. Come, join our NVLink design team and help build the real-time, cost-effective computing platform driving our success in this exciting and quickly growing field.
Legal Program Manager, Compute & Infrastructure AnthropicLegal Program Manager, Compute & InfrastructureSan Francisco, CA10+ years in contract management, legal operations, or procurement operations, including hands-on experience implementing or administering a contract lifecycle management (CLM) platform; experience supporting infrastructure, datacenter, hardware, construction, or telecom/network procurement and vendor contracting. The ability to thrive as a member of cross-functional teams building frontier technologies, with a desire to develop a deep understanding of the Compute & Infrastructure teams and the datacenter, hardware, and network systems they build and operate.
Senior Infrastructure Automation Engineer, Compute Platform NvidiaSenior Infrastructure Automation Engineer, Compute PlatformSanta Clara, CAWorking alongside our LSF internals engineer to encode hard-won scheduler knowledge into templates and policy, so that expertise lives in the repository instead of in one person's head. What you'll be doing: Designing and owning the configuration schema for LSF cell deployment, so that a policy change is written once, reviewed, tested, and applied identically everywhere it belongs.
Portal & Infrastructure Tech Lead Faraday FuturePortal & Infrastructure Tech LeadFremont, CA$155,000–$180,000 / yearThe ideal candidate is a seasoned platform engineer who has built systems used by large external developer communities, understands what makes an API a pleasure or a pain to work with, and has the leadership range to grow a high-performing engineering team while remaining deeply technically engaged. Lead the design and engineering of the platform's core developer-facing products: the robotics APIs (covering motion control, perception, sensor fusion, autonomous decision-making, and inter-device communication) and the multi-language SDKs (Python, Node.js, Go, Java, and others as needed).
NewKey Account Manager - Infrastructure (2517) HellermannTytonKey Account Manager - Infrastructure (2517)San Francisco, CARemoteThe Key Account Manager - Infrastructure will play a pivotal role in driving strategic development and sales growth within the Data Center market, which includes but is not limited to Construction Companies that design and build Data Centers, Data Center owners and their Engineers. Excellent verbal and written communication skills, including the ability to recognize and customize communications to different audiences, including utilizing diverse information from a variety of sources to present the HellermannTyton value proposition in an effective manner.
Staff Software Developer - IVI Test Infrastructure Ford Motor CompanyStaff Software Developer - IVI Test InfrastructurePalo Alto, CAThis is a hands-on IC lead role where you'll define infrastructure, guide adoption of best practices, streamline developer tools, and help shape the future of how Ford delivers reliable, secure, and high-quality EV software. In this role, you will design, build, and scale the test infrastructure that developers, integration engineers, and QA teams rely on to validate software across vehicles, mobile, and cloud.
Staff Infrastructure Engineer ReplitStaff Infrastructure EngineerFoster City, CaliforniaYou will design robust monitoring solutions, automate operational tasks, and continuously improve our infrastructure's reliability, all while mentoring and educating the broader engineering team to make reliability a core value at Replit. Drive Cross-Company Improvements: Partner directly with service owners across Replit to understand their pain points, and collaborate on implementing build/test/deploy enhancements within their specific services.