Senior Site Reliability Engineer- San Francisco, CA, The US KodySenior Site Reliability Engineer- San Francisco, CA, The USSan Francisco, CASenior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America.
Senior CockroachDB Database Engineer / Site Reliability Engineer (SRE) Resource Logistics, Inc.Senior CockroachDB Database Engineer / Site Reliability Engineer (SRE)Sunnyvale, CAWe are seeking a highly skilled CockroachDB Database Engineer with strong Site Reliability Engineering (SRE) experience to design, implement, manage, and optimize large-scale distributed database platforms. The role requires close collaboration with development, infrastructure, and platform engineering teams to ensure highly available, resilient, and scalable database services.
Senior Inference Reliability Engineer ParasailSenior Inference Reliability EngineerSan Mateo5+ years of experience in production engineering, site reliability engineering, infrastructure engineering, distributed systems, ML infrastructure, database reliability, or performance engineering. You may come from GPU infrastructure, ML platforms, databases, streaming systems, search, low-latency services, distributed storage, or large-scale production engineering.
NewSystems Quality And Reliability Engineer - LPU NvidiaSystems Quality And Reliability Engineer - LPUSanta Clara, CAManage operational perf of FA at CMs, ensuring partner achieve key perf indicators including FA cycle times, fault duplication rates and fault isolation rates. The base salary range is 136,000 USD - 218,500 USD for Level 3, and 168,000 USD - 264,500 USD for Level 4. You will also be eligible for equity and benefits.
Lead Reliability Engineer SabiLead Reliability EngineerSan Francisco, CaliforniaPartner with the EE lead on EMC and ESD; with audio and camera leads on component-level reliability; with the product design lead on drop, ingress, and thermal robustness; with the firmware lead on field telemetry and diagnostics. Stand up our test capability, internal where it makes sense and external lab partnerships where it doesn't, and own the test fixtures, sample plans, and pass/fail criteria that gate every release.
NewSenior Cell Validation & Reliability Engineer Peak EnergySenior Cell Validation & Reliability EngineerBurlingame, CA$150,000–$190,000 / yearFirst, validate the electrical and mechanical integration of supplier cells into Peak Energy's energy storage systems — moving beyond verification of the cell specification to quantifying how population variance (cell-to-cell and batch-to-batch) impacts system-level performance. Move validation from spec verification to population statistics: design cell-to-cell and batch-to-batch variation studies and quantify how variance propagates to system-level performance — balancing burden, thermal gradients, mechanical preload windows, and energy/power delivery.
Site Reliability Engineer SpecterSite Reliability EngineerSan Francisco, CaliforniaYou'll drive reliability across our sensor fleet — triaging issues in the field, building the systems that prevent them from recurring, and owning the observability that keeps us ahead of problems as we scale. We're hiring a Site Reliability Engineer to own the operational health of our connected sensor platform — spanning a live fleet of edge hardware deployed at customer sites and the cloud infrastructure behind it.
NewPlatform Site Reliability Engineer SpecterPlatform Site Reliability EngineerSan Francisco, CaliforniaYou’ll operate and improve our Kubernetes-based infrastructure, manage cloud resources through Terraform, strengthen observability and incident response, and build the systems that allow our engineering teams to deploy safely and move quickly. Design and improve observability into system health, functionality, and performance through logging, metrics, tracing, dashboards, and alerting across Kubernetes workloads and AWS infrastructure.
Senior Platform & Reliability Engineer OpenArt AISenior Platform & Reliability EngineerSan Francisco, CaliforniaYou’ll partner closely with product engineers to evolve the platform that powers OpenArt, contributing to key decisions around infrastructure architecture, improving multi-provider AI reliability, and helping us scale systems to millions of users—while raising the overall engineering bar. You’ll work across cloud infrastructure, distributed systems, backend services, and developer tooling, making pragmatic decisions that balance product velocity, system reliability, and cost efficiency—in a fast-moving, AI-native environment.
NewReliability Engineer Macpower Digital Assets Edge Private LimitedReliability EngineerSan Leandro, CA$65–$70 / hourEssential Functions and Responsibilities: Conduct root cause analysis (RCA) on high-cost failures, high-downtime, and repetitive issues to identify engineering solutions. Job Summary: Enhance equipment reliability and drive plant-wide continuous improvement across maintenance, operations, safety, quality, and training.
Sr. Reliability Engineer MindlanceSr. Reliability EngineerMilpitas, CA$44.64–$47.48 / hourMinimum of 5 years of reliability engineering experience or experience performing tests to collect experimental data and performing statistical analyses to interpret results. Manage reliability projects by implementing tests, analyzing data, assessing reliability risks, and reporting accurate information to stakeholders.
Linux Private Cloud Site Reliability Engineer Eitacies IncLinux Private Cloud Site Reliability EngineerSanta Clara, CARemoteFull timeWe are seeking experienced Linux Site Reliability Engineers to support large-scale hybrid infrastructure environments across cloud and private data centers. Candidates should have deep Linux administration experience combined with infrastructure automation and production troubleshooting.
Platform Site Reliability Engineer Eitacies IncPlatform Site Reliability EngineerSanta Clara, CARemoteFull timeWe are hiring experienced Platform Site Reliability Engineers to build internal platform services that improve developer productivity, deployment automation, CI/CD workflows, and platform reliability. This position is ideal for engineers who enjoy building reusable platform capabilities rather than supporting individual applications.
Kubernetes Site Reliability Engineer Eitacies IncKubernetes Site Reliability EngineerSanta Clara, CARemoteFull timeThis role focuses on building reliable, scalable, and automated infrastructure using Kubernetes, AWS, Terraform, Linux, and modern DevOps practices. We are hiring multiple Site Reliability Engineers to support enterprise Kubernetes platforms running in secure cloud environments.
Principal Site Reliability Engineer Cloud Identity & Trust SPIFFE/SPIRE ESRhealthcare and EXEC STAFF RECRUITERSPrincipal Site Reliability Engineer Cloud Identity & Trust SPIFFE/SPIRESan Jose, CaliforniaProficiency in operating and supporting cloud-based services using IaC (infrastructure as code, Terraform Proven experience as a systems administrator or service reliability engineer Experience with CI/CD processes and source control mechanisms (GitHub). Experience level: Mid-senior Experience required: 10 Years Education level: Bachelors degree Job function: Information Technology Industry: Information Technology and Services Pay rate: $60 per hour Total position: 1 Relocation assistance: No Visa sponsorship eligibility: No.
Senior/Staff Power Electronics Hardware Reliability Engineer ZiplineSenior/Staff Power Electronics Hardware Reliability EngineerSan Francisco, CA$140,000–$230,000 / yearDeep experience developing and validating safety-critical high-voltage and high-power electrical systems such as grid-connected converters, industrial power supplies, EV charging systems, inverters, high band gap semiconductors, renewable-energy systems, aerospace power systems, or similar equipment. Our customers include the world's largest and most prominent healthcare systems, governments, retailers, restaurants and global businesses who rely on us to save lives, reduce emissions, increase economic opportunity, and provide delivery from point A to point B as fast as possible.
Reliability Engineer Bright Vision TechnologiesReliability EngineerMilpitas, CARemoteFull timeThe ideal candidate will combine deep systems knowledge with strong programming skills, a measurement-driven mindset, and the discipline to design, automate, and operate complex services so that reliability becomes a first-class engineering deliverable rather than a reactive concern. We are seeking an experienced Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production.
Reliability Engineer Intel Corp.Reliability EngineerSanta Clara, CA$122,440–$232,190 / yearResponsibilities: Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) for compute, memory, storage, network, power, and cooling subsystems. Mission: Define and own the pod-level reliability specifications that ensure the availability, resilience, and serviceability of a large-scale data center across hardware, thermal, and operational dimensions.
Failure Analysis and Reliability Engineer Advanced Micro Devices, IncFailure Analysis and Reliability EngineerSan Jose, CaliforniaThis role offers the opportunity to build deep hands-on failure analysis expertise, work with advanced products and analytical tools, and contribute directly to product quality, reliability, and customer success. This is a hands-on engineering position suited for someone who enjoys solving complex technical problems, working across teams, and driving issues from initial investigation through root-cause conclusion.
Senior Reliability Engineer Eight SleepSenior Reliability EngineerSan Francisco, CaliforniaWe're looking for a hands-on, detail-oriented engineer who is passionate about identifying failure risks early in the design cycle, driving cross-functional collaboration, and delivering hardware that performs reliably in the harshest real-world environments. Leverage data analytics tools (e.g., Python, JMP, MATLAB) to automate testing workflows, perform reliability modeling (e.g., Weibull, Monte Carlo), and extract actionable insights from test data.