Senior Site Reliability Engineer- Sunnyvale, CA, the US KodySenior Site Reliability Engineer- Sunnyvale, CA, the USSunnyvale, CAYou will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America. Data-driven mindset with the ability to leverage metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements.
Senior Site Reliability Engineer- Palo Alto, the US KodySenior Site Reliability Engineer- Palo Alto, the USPalo Alto, CAYou will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America. Data-driven mindset with the ability to leverage metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements.
Senior Site Reliability Engineer- San Francisco, CA, the US KodySenior Site Reliability Engineer- San Francisco, CA, the USSan Francisco, CAYou will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America. Data-driven mindset with the ability to leverage metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements.
Staff TDI Site Reliability Engineer, Okta Federal OktaStaff TDI Site Reliability Engineer, Okta FederalSan Francisco, CA$174,000–$239,000 / yearNotice for New York City Applicants & Employees: Okta may use Automated Employment Decision Tools (AEDT), as defined by New York City Local Law 144, that use artificial intelligence, machine learning, or other automated processes to assist in our recruitment and hiring process. The annual base salary range for this position for candidates located in California (excluding San Francisco Bay Area), Colorado, Illinois, New York, and Washington is between: $174,000—$239,000 USD The Okta Experience .
Reliability Engineer, Mechanical Systems, NA Vantage Data CentersReliability Engineer, Mechanical Systems, NASanta Clara, CaliforniaFor each of the major systems Electrical, Mechanical, and Controls, the Reliability Engineering team is responsible for ensuring success in the commissioning stages of new construction, evaluating and improving the reliability and performance of existing critical infrastructure, sustaining equipment operational availability through maintenance program design, providing ongoing technical support to the Site Operations Teams, as well as providing systems reliability and maintainability feedback to the Design Engineering teams for future design considerations. Developing and operating across North America, EMEA and Asia Pacific, Vantage has evolved data center design in innovative ways to deliver dramatic gains in reliability, efficiency and sustainability in flexible environments that can scale as quickly as the market demands.
Sr Staff Reliability Engineer Aviso, LLCSr Staff Reliability EngineerFremont, CAFull timeApply Advanced Methodologies: Utilize reliability engineering tools such as Weibull analysis, Reliability Modeling, Fault Tree Analysis, Accelerated Life Testing (ALT), HALT/HASS, and structured root cause analysis to drive data-based design decisions. Define Requirements: Establish reliability requirements, reliability growth plans, and demonstration strategies for NPI (New Product Introduction) programs and sustaining design changes; devise test protocols and author comprehensive test reports.
Sr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid) CrowdStrikeSr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid)Sunnyvale, CA$140,000–$215,000 / yearSeveral key feature additions and critical re‑architecture initiatives are underway: deploying core services to additional cloud providers, modularizing into reusable components, improving core libraries and frameworks, maturing observability tooling (tracing, profiling, alerting, SLOs), and automating away manual toil across all of the above. You will shape architectural choices that affect every feature development team and provide shared architectural components leveraged throughout the Falcon Platform - including Unified Search, Protobuf libraries, and other shared‑tier infrastructure.
NewSenior Site Reliability Engineer : 26-02282 Akraya, Inc.Senior Site Reliability Engineer : 26-02282San Jose, CARemote$55–$60 / hourMost recently, we were recognized Stevie Employer of the Year 2025, SIA Best Staffing Firm to work for 2025, Inc 5000 Best Workspaces in US (2025 & 2024) and Glassdoor's Best Places to Work (2023 & 2022)! This role seeks a seasoned Site Reliability Engineer to develop and manage critical platform services for secure, high-compliance government environments.
Senior Site Reliability Engineer DevelocitySenior Site Reliability EngineerSan Francisco, CARemote$180,000–$205,000 / yearDevelocity helps software teams achieve delivery excellence through deep observability, build and test acceleration, and AI-powered intelligence across the entire toolchain – with current support for Gradle Build Tool, Apache Maven, sbt, npm, and Python. You'll be responsible for the reliability, performance, and availability of Develocity instances serving paying customers, open-source projects, and public-facing services, plus supporting infrastructure like artifact registries.
Robotics Hardware Reliability Engineer Dyna RoboticsRobotics Hardware Reliability EngineerRedwood City, CaliforniaBench-test suspect actuators and electronics, rework and repair (soldering, harness and connector work), swap and re-commission components, and run physical validation with calipers, test fixtures, and boot and power-cycle rigs you build yourself. Working knowledge of robot control: impedance and torque-controlled actuators, force-control parameters, and controller state lifecycles (boot, e-stop, reconnect).
Product Quality And Reliability Engineer IV Applied MaterialsProduct Quality And Reliability Engineer IVSanta Clara, CA$133,500–$183,500 / yearIf you would like to contact us regarding accessibility of our website or need assistance completing the application process, please contact us via e-mail at Accommodations_Program@amat.com, or by calling our HR Direct Help Line at 877-612-7547, option 1, and following the prompts to speak to an HR Advisor. Drive DfQR for key programs to develop FMECA/technical risks, apply BKMs/lessons, learn to reduce risks, and validate solutions by developing and tracking CRAMS/ERAMS test plans.
Site Reliability Engineer TELCORSite Reliability EngineerSan Francisco, NebraskaCopy and paste the following link into your browser to learn more about TELCOR and what it means for TELCOR to be a certified Great Place To Work®: https://www.greatplacetowork.com/certified-company/7054288. This role will also design and operate resilient systems across cloud and containerized environments, as well as manage production infrastructure and deployment workflows across environments.
Site Reliability Engineer VantageScoreSite Reliability EngineerSan Francisco, CaliforniaCollaborate with IT & Info-Sec SMEs on AWS IAM roles and policies, VPC configurations, Security Groups, CloudTrail, Config, and GuardDuty to ensure least-privilege access and auditability. Define and track SLOs/SLAs for internal and external APIs; implement alerting and dashboards using observability tooling (e.g., CloudWatch, Datadog, Grafana).
Senior Silicon Photonics TD Reliability Engineer Intel Corp.Senior Silicon Photonics TD Reliability EngineerSanta Clara, CA$141,910–$232,190 / yearBusiness group: Intel Foundry strives to make every facet of semiconductor manufacturing state-of-the-art while delighting our customers -- from delivering cutting-edge silicon process and packaging technology leadership for the AI era, enabling our customers to design leadership products, global manufacturing scale and supply chain, through the continuous yield improvements to advanced packaging all the way to final test and assembly. Your new team, (SiPh TD QnR), is at the forefront of silicon photonics technology development, working closely and collaborating with the SiPh fab process integration, Integrated Photonics Solutions (IPS) R&D design team, Foundry Q&R teams and other organizations to ensure reliability for internal and external SiPh device technologies, process certification, and product qualifications.
NewMember of Technical Staff, Site Reliability Engineer InferactMember of Technical Staff, Site Reliability EngineerSan Francisco, CaliforniaYou'll work across engineering and infrastructure to define SLOs, improve monitoring and alerting, strengthen incident response, drive post-mortems, and reduce operational risk before it reaches users. We're looking for a Site Reliability Engineer to help make vLLM-powered inference systems reliable, observable, and operationally simple at production scale.
Senior Software Engineer, Infrastructure Nectar SocialSenior Software Engineer, InfrastructurePalo Alto, CaliforniaComprehensive stipends: $1,000/month housing stipend for living near the office, $50/month mobile stipend, $50/month internet stipend for remote employees, and commuter benefits (train stipend or Palo Alto parking permit reimbursement). We run high-volume data ingestion pipelines and real-time AI agents on top of a fast-growing customer base, and we need a seasoned SRE to help us scale these systems safely and keep them running flawlessly.
Senior Site Reliability Engineer AirbyteSenior Site Reliability EngineerSan Francisco, CaliforniaYou'll be the infrastructure and reliability engineer on the Data Replication team - a full-stack product team running over 3 million sync jobs a week powering thousands of data use cases across multiple regions and clouds. We give agents fast, accurate, authenticated access to business data across hundreds of sources, so they can discover the entities that matter, reason over real-time context, and take action in the systems they read from, not just observe them.
Senior Site Reliability Engineer- Sunnyvale, CA, The US KodySenior Site Reliability Engineer- Sunnyvale, CA, The USSunnyvale, CASenior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America.
Reliability Engineer (Hardware) LightmatterReliability Engineer (Hardware)Mountain View, CA$142,000–$200,000 / yearYou will enable our products to meet or exceed reliability requirements throughout their lifecycle by developing robust test protocols, analyzing failures, and working with cross-functional teams to drive improvements in design and manufacturing processes. Perform detailed root cause analysis of device failures, and collaborate with cross-functional teams to resolve reliability issues, refine product designs, and improve process robustness.
Senior Site Reliability Engineer- Palo Alto, The US KodySenior Site Reliability Engineer- Palo Alto, The USPalo Alto, CASenior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America.