Reliability Engineer/FA Tanisha SystemsReliability Engineer/FASanta Clara, CAFull timeWhat you will be doing: Work in the Board Level Reliability lab environment and set up functional test hardware and software for various NV products, including large server systems, and perform functional validation for GPU/Tegra products. Conduct advanced failure analysis using techniques such as CSAM (C-Scan Acoustic Microscopy) and X-ray imaging to detect internal defects, delamination, voids, and structural issues.
Staff Site Reliability Engineer Circle Internet FinancialStaff Site Reliability EngineerSan Francisco, CaliforniaRemote$195,000–$257,500 / yearWorking closely with engineering, protocol, product, and security teams, you’ll build AI-powered tooling and automation that improves operational excellence, accelerates developer productivity, and enables the rapid delivery of new blockchain capabilities. All the requirements of a Senior Site Reliability Engineer and: 6+ years of experience in SRE, DevOps, or Infrastructure Engineering, ideally within a cloud-native or blockchain environment; Proven track record of technical leadership in complex, distributed systems architecture and design.
Senior Site Reliability Engineer FiservSenior Site Reliability EngineerSunnyvale, CaliforniaFor incentive eligible associates, the successful candidate is eligible for an annual incentive opportunity which may be delivered as a mix of cash bonus and equity awards in the Company’s sole discretion. You will partner with cross-functional teams to improve reliability, automate operations, and drive continuous improvement across our cloud-native environments.
Senior Site Reliability Engineer OutSystemsSenior Site Reliability EngineerMenlo Park, CaliforniaAs an SRE at OutSystems here are your key responsibilities and duties: Lead and onboard services and teams to the reliability tenets; Establish and maintain Service Level Objectives (SLOs) and Service Level Agreements (SLAs); Design and implement scalable, reliable, and secure infrastructure, while ensuring cloud-native best practices; Collaborate with software development teams to ensure systems are resilient (observable, fault-tolerant, recoverable, scalable) and performant; Implement monitoring, alerting, logging, and tracing solutions to detect and respond to incidents; Lead incident response efforts, ensuring quick resolution and minimal downtime, and conduct RCA/post-mortems; Automate every operational task, with a special focus on fast incident detection & recovery; Programming in Python supported by Gen AI tooling to accelerate development of mission critical automation and tools. (CKA, CKAD, CKS certifications are valued); Experience with automation and Infrastructure as Code (IaC) tools, such as AWS CloudFormation, Terraform, Puppet, Chef, Spacelift, etc; Experience with Python, Go, Bash/Shell scripting, or other automation tools/languages; Familiarity with AWS services like EC2, RDS, ELB, CloudFront, Lambda, etc; Proficiency in monitoring and troubleshooting complex distributed systems; Experience with Grafana, ELK stack, Prometheus, or others; Strong understanding of designing resilient and fault-tolerant systems; Expertise in debugging complex distributed systems.
Customer Reliability Engineer Andromeda ClusterCustomer Reliability EngineerSan Francisco, CaliforniaRemoteWork GPU failures: driver and device-plugin issues, XID errors, thermal throttling, nodes that need cordoning or draining, jobs failing across multiple nodes. Dig into Kubernetes problems like pods stuck pending or crash-looping, node conditions, scheduling failures, resource limits.
Site Reliability Engineer Sustainable TalentSite Reliability EngineerSanta Clara, CA$65–$85 / hourThis group works with various other groups within NVIDIA Software such as Graphics Processors, Mobile Processors, Deep Learning, Artificial Intelligence and Driverless Cars to cater to their infrastructure needs. The cloud hosts a heterogeneous mix of machines and devices with various operating systems (Windows/Linux/Android), a multitude of hardware platforms both NVIDIA GPUs and Tegra Processors.
Sr. IAM Site Reliability Engineer Pure Storage Inc.Sr. IAM Site Reliability EngineerSanta Clara, CA$186,000–$279,000 / yearClassification of protected categories is as follows: A "disabled veteran" is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service-connected disability. Change Governance & Cross-Functional Collaboration: Partner across technical and business teams to lead change management risk assessments, deploy new identity capabilities, and maintain actionable runbooks and architecture documentation to ensure transparent, repeatable operations.
Site Reliability Engineer Recruiting From ScratchSite Reliability EngineerSan Francisco, CaliforniaThis organization is focused on developing advanced medical language models that streamline processes and improve efficiency in the healthcare sector. We're representing a dynamic company at the forefront of healthcare innovation, leveraging AI to automate operations and enhance patient care.
Lead Database Reliability Engineer - 11606 Coupa SoftwareLead Database Reliability Engineer - 11606San Francisco, CARemote$142,000–$198,667 / yearBy submitting your application, you acknowledge that you have read Coupa's Privacy Policy and understand that Coupa receives/collects your application, including your personal data, for the purposes of managing Coupa's ongoing recruitment and placement activities, including for employment purposes in the event of a successful application and for notification of future job opportunities if you did not succeed the first time. Collaborate effectively across cross-functional teams, mentor junior database engineers, stay current on emerging database technologies and best practices, and remain flexible to support global teams across multiple time zones.
Senior Site Reliability Engineer Hyperbolic LabsSenior Site Reliability EngineerSan Francisco, CaliforniaYou'll be responsible for defining and maintaining service level objectives for job success rates, building robust incident response systems, managing capacity across our distributed GPU network, and implementing secure rollout and rollback mechanisms that keep our platform running smoothly 24/7. Architected, deployed, and managed large-scale Kubernetes environments, including cluster administration, container orchestration, autoscaling, service discovery, and high-availability infrastructure to ensure reliability and scalability of mission-critical systems.
Software Reliability Engineer NuroSoftware Reliability EngineerMountain View, CA$145,830–$219,000 / yearWith years of real-world deployment experience and a flexible, partner-led business model, Nuro is working toward a future where millions of autonomous vehicles powered by our technology help make everyday life safer, easier, and more connected. Powered by the Nuro Driver, our universal autonomy platform enables the global mobility ecosystem to deploy autonomy at scale, from robotaxis and logistics fleets to personal vehicles.
Hardware Reliability Engineer SkipHardware Reliability EngineerSan Francisco, CAOwn the reliability testing program for Skip's wearable devices -- design test plans, define pass/fail criteria, execute testing across EVT/DVT/PVT builds, and track results through to closure. with Arc'teryx) we are uniquely positioned to launch the first commercially successful wearable robotic device, the MO/GO, develop a platform to launch future Movewear products and transform millions of lives in the coming years.
Senior Site Reliability Engineer - Core Cloud Platform LambdaSenior Site Reliability Engineer - Core Cloud PlatformSan Francisco, CaliforniaOur investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove. *Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
Advanced Packaging Reliability Engineer OpenAIAdvanced Packaging Reliability EngineerSan Francisco, CaliforniaFor unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In-depth knowledge of advanced packaging architectures, including 2.5D and 3.5D integration, interposers, embedded bridges, chiplets, large package substrates, HBM integration, redistribution layers, and package-level power delivery.
Staff Security Reliability Engineer OpenAIStaff Security Reliability EngineerSan Francisco, CAFor unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. We are looking for an experienced Site Reliability Engineer working on security infrastructure to design, build, and operate reliable, secure, and scalable infrastructure that underpins identity, access, endpoint, and shared platform services across the company.
Senior Site Reliability Engineer Autodesk Inc.Senior Site Reliability EngineerSan Francisco, CA$117,000–$209,330 / yearThe ideal candidate has deep experience operating production systems at scale, an automation-first mindset, and the ability to improve reliability through engineering practices such as SLOs/SLIs, production readiness, incident management, observability, resilience testing, and toil reduction. As part of a new SRE team supporting Autodesk GovCloud, you will have a unique opportunity to help shape how Autodesk deploys, runs, and improves production services in restricted cloud environments.
Sr. Site Reliability Engineer PayNearMe, Inc.Sr. Site Reliability EngineerSanta Clara, CA$180,000–$200,000 / yearOur single platform handles it all: cards, ACH, digital wallets such as PayPal, Venmo, Cash App Pay, Apple Pay and Google Pay, and even cash at more than 62,000 retail locations nationwide. Tool Development: Ability to write and update tools to support infrastructure and application management, demonstrating the principle that "SRE is what happens when you ask a software engineer to design an operations team.
Site Reliability Engineer (SRE) Recruiting From ScratchSite Reliability Engineer (SRE)San Francisco, CaliforniaAs demand for AI infrastructure accelerates, they're investing heavily in reliability engineering to build the automation, observability, and platform infrastructure that powers their multi-cloud GPU marketplace at scale. 3–10 years of experience in Site Reliability Engineering, Production Engineering, Infrastructure Engineering, or Platform Engineering.
Site Reliability Engineer Forward NetworksSite Reliability EngineerSanta Clara, CA$230,000–$250,000 / yearBacked by world-class investors including Andreessen Horowitz, Goldman Sachs, MSD Partners, and Threshold Ventures, Forward offers a people-centric, innovative culture where brilliant minds are shaping the future of network reliability, security, and AI-ready operations. As our first or early SRE hire you will be building the reliability engineering function at Forward — defining how we think about availability, observability, incident response, and operational excellence across a complex, distributed SaaS platform.
Senior Site Reliability Engineer FiveTranSenior Site Reliability EngineerOakland, CA$174,915–$209,906.50 / yearFivetran promotes diversity, equity, inclusion & belonging through attracting, recruiting, developing, and retaining a diverse workforce, not only because it is the right thing to do, but because it helps us build a world-class company to better serve our customers, our people and our communities. Together, we're delivering the data infrastructure layer that helps organizations move, transform, and trust their data - from the moment data moves, through every transformation, to the context teams and AI systems rely on.