NewHardware Reliability Engineer Ouraring IncHardware Reliability EngineerSan Francisco, CA$172,000–$203,000 / yearBuild test-to-field correlation for failure modes identified in reliability testing, ensuring lab test methods accurately predict real-world product performance and continuously refining test plans based on field learnings. Perform Design Failure Mode and Effects Analysis (DFMEA) to proactively identify potential risks and failure modes early in the design process, driving mitigation actions ahead of validation phases.
Reliability Engineer, Control Systems, NA Vantage Data Centers Management Co LLCReliability Engineer, Control Systems, NASanta Clara, CAFor each of the major systems Electrical, Mechanical, and Controls, the Reliability Engineering team is responsible for ensuring success in the commissioning stages of new construction, evaluating and improving the reliability and performance of existing critical infrastructure, sustaining equipment operational availability through maintenance program design, providing ongoing technical support to the Site Operations Teams, as well as providing systems reliability and maintainability feedback to the Design Engineering teams for future design considerations. Developing and operating across six markets in North America and five markets in Europe, Vantage has evolved data center design in innovative ways to deliver dramatic gains in reliability, efficiency and sustainability in flexible environments that can scale as quickly as the market demands.
Site Reliability Engineer II Maplebear IncSite Reliability Engineer IICARemote$160,000–$169,000 / yearWith members from diverse backgrounds and experiences, SRE fosters a supportive and risk-tolerant environment where individuals can think big, take on meaningful projects, and grow with guidance and mentorship. The Site Reliability Engineering (SRE) team integrates software and systems engineering to design and manage large-scale, distributed, and fault-tolerant systems.
Senior Site Reliability Engineer - Core Cloud Platform Lambda IncSenior Site Reliability Engineer - Core Cloud PlatformSan Jose, CAOur investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove. Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda's designated work from home day is currently Tuesday.
Module Reliability Engineer - Imaging and Sensing Apple IncModule Reliability Engineer - Imaging and SensingCupertino, CAscanning electron microscopy, CT scanning, FTIR,…) Excellent written and verbal communication skills with a strong attention to detail Ability to travel internationally without restrictions (up to 10%)A graduate degree in Materials Science, Physics, Engineering or equivalent field 3-5 years of experience in an engineering role which involved hands on work Possess extraordinary problem solving abilities Have experience with the design and manufacture of assemblies that convert digital signals into physical signals (voltage, light, vibration) and vice versa Are adept in statistical data analysis and are capable of clear reporting Ability to handle multiple projects at once including running activities remotely at third party vendor sites overseas Ability to complete tests and resolve problems in an unattended manner Flexibility to accept new tasks and go beyond ones comfort zone to pursue the solution to new and unexpected problems Dynamic person with a can do attitude and a dream to work with a great team and product. Drive & update FMEA Establish specifications for component/module reliability Set up reliability test infrastructure at Apple suppliers Perform validation test development Generate reliability test plans and drive test execution at Apple suppliers Analyze and summarize reliability data Drive failure analysis and mitigations Provide risk assessmentsA Bachelor's degree in Materials Science, Physics, Engineering or equivalent field with 3+ years of industry experience Understand the physics behind the mechanical & electrical properties of materials Have exposure to failure analysis techniques as well as Materials characterization tools (e.g.
Site Reliability Engineer, Enterprise Technology Services Apple IncSite Reliability Engineer, Enterprise Technology ServicesSunnyvale, CALead Operational Excellence & Incident Management: Drive comprehensive operational excellence through advanced observability (tracing, logging, metrics, alerting) and next-generation telemetry, leveraging Machine Learning for anomaly detection and exploring GenAI for alert engineering. Promote a strong DevOps culture and provide technical insights through log analysis and system debugging.5+ years of experience in Site Reliability Engineering with a strong focus on building, scaling, and operating large-scale distributed platform services, and Java.
Senior Reliability Engineer Also IncSenior Reliability EngineerPalo Alto, CAWe're a passionate team of builders, dreamers, doers and innovators, focused on creating entirely new (not to mention, innovative and delightful) vertically integrated, small EVs designed to meet the global mobility challenges of today and tomorrow. Develop new reliability tests procedures and specifications such as highly accelerated life testing, environmental testing, and other reliability tests to demonstrate reliability.
Staff Field Reliability Engineer Also IncStaff Field Reliability EngineerPalo Alto, CA$200,000–$240,000 / yearTwo or more years of industry experience in a reliability engineering roleTechnical knowledge of one or more aspects related to reliability, including: PCBA and electronic components, battery and power electronics, drive units, and chassis systemsUnderstanding of reliability statistics, life data analysis, reliability growthExperience with structured problem solving like FRACA or 8D, with root cause analysis (RCA) tools like: 5 Whys, Fishbone, Fault Tree, FMEAWorking knowledge of a coding language, preferably Python. Working knowledge of statistical software for reliability, such as JMP, the ReliaSoft suite, or MATLABWorking knowledge of failure analysis techniques such as SEM, CT, X-Ray, EDS, TDRExperience with deploying reliability testing guidelines and inventing new ways of testing.
Hardware Reliability Engineer - Apple Vision Products Apple IncHardware Reliability Engineer - Apple Vision ProductsCupertino, CAGuide design and interact with diverse groups to improve product reliability Statistical data analysis to evaluate designs and provide design risk assessments Present test plans, results, and risk assessments in a confident and positive manner to management and executive teams Develop new reliability tests procedures and specifications Prepare concise and detailed test plans and analyze test results Research new technologies to understand unique failure mechanisms Bachelor's Degree in Mechanical Engineering, Materials Science, Physics or an equivalent field 3+ years of experience in a related field or industry Understand the physics behind the mechanical & electrical properties of materials Have exposure to failure analysis techniques and the ability to use failure analysis methodology to derive a root cause of failure Ability and willingness to travel internationally, as needed (up to 10 percent)M.S. or PhD in Mechanical Engineering, Materials Science, Physics or an equivalent field Statistical experience such as Weibull, JMP, or familiarity with accelerated test models Solid understanding of Failure Modes and Effects Analysis (FMEA) and Design of Experiments (DOE) Excellent communication skills, both written and verbal Ability to manage multiple projects simultaneously Strong attention to details and curiosity of how technologies work Collaborative attitude and ability to work cross-functionally effectively A proactive, can-do attitude, with a passion for working alongside an amazing team and innovative products. As a System Reliability Engineer, youll be responsible for identifying high-risk failure modes early in the design process and working closely with design engineering teams to reduce those risks.
Site Reliability Engineer Amiri RecruitingSite Reliability EngineerMountain View, CaliforniaBuild, maintain, and optimize Kubernetes clusters (including GPU-backed clusters). Administer and improve our platform architecture and apply general security best practices across the stack.
Staff Site Reliability Engineer- Developer Platform Rivian and Volkswagen Group TechnologiesStaff Site Reliability Engineer- Developer PlatformPalo Alto, CaliforniaRivian and Volkswagen Group Technologies may share your Candidate Personal Data with (i) internal personnel who have a need to know such information in order to perform their duties, including individuals on our People Team, Finance, Legal, and the team(s) with the position(s) for which you are applying; (ii) Rivian and Volkswagen Group Technologies affiliates; and (iii) Rivian and Volkswagen Group Technologies’ service providers, including providers of background checks, staffing services, and cloud services. Rivian and Volkswagen Group Technologies may use your Candidate Personal Data for the purposes of (i) tracking interactions with our recruiting system; (ii) carrying out, analyzing and improving our application and recruitment process, including assessing you and your application and conducting employment, background and reference checks; (iii) establishing an employment relationship or entering into an employment contract with you; (iv) complying with our legal, regulatory and corporate governance obligations; (v) record keeping; (vi) ensuring network and information security and preventing fraud; and (vii) as otherwise required or permitted by applicable law.
Systems Reliability Engineer (SRE) ClaryoSystems Reliability Engineer (SRE)San Francisco, CaliforniaExperience working with networking in constrained or distributed environments (e.g., VPNs, secure tunnels, on-site networking). Our platform runs across distributed infrastructure—connecting cloud services, on-site compute, and live video/data pipelines inside warehouses.
Integration Reliability Engineer ClaryoIntegration Reliability EngineerSan Francisco, CaliforniaExperience working with networking in constrained or distributed environments (e.g., VPNs, secure tunnels, on-site networking). Our platform runs across distributed infrastructure—connecting cloud services, on-site compute, and live video/data pipelines inside warehouses.
Site Reliability Engineer BaseTenSite Reliability EngineerSan Francisco, CAYou'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production.
NewSenior Site Reliability Engineer Andromeda ClusterSenior Site Reliability EngineerSan Francisco, CaliforniaIncident Management: Proven track record leading incident response for complex distributed systems where the failure could be in hardware, firmware, networking, drivers, orchestration, or application code and you need to narrow it down fast. Observability & Monitoring: Hands-on experience building monitoring and alerting for GPU infrastructure, not just Prometheus/Grafana basics, but GPU-specific telemetry (DCGM, nvidia-smi, fabric manager metrics) integrated into actionable dashboards.
Senior Site Reliability Engineer, Apple Data Platform Compute SRE Apple IncSenior Site Reliability Engineer, Apple Data Platform Compute SRECupertino, CAThis role focuses on managing bare-metal and cloud based infrastructure, levering and extending our infrastructure-as-code based tooling, analyzing and optimizing performance, helping to plan and execute long term fleet management logistics, capacity planning, and ultimately maintaining operational excellence across distributed data platforms that power analytics across Apple. We partner with both peer SRE teams and several of our world-class software and product engineering teams to support infrastructure reliability, multi-year parallel migrations for Apple properties, as well as the automation, tooling, incident, and process management necessary to ensure smooth 24x7 operations for ADP customers.
Senior SSD Reliability Engineer - Contract SK hynix memory solutions America Inc.Senior SSD Reliability Engineer - ContractSan Jose, CAAs a test engineer, you will help design, develop, automate, and execute test plans to validate the functionality, endurance, and reliability of our enterprise data center SSD products. As a global leader in DRAM and NAND flash technologies, we drive the evolution of advancing mobile technology, empowering cloud computing, and pioneering future technologies.
iPhone System Reliability Engineer AppleiPhone System Reliability EngineerCupertino, CAFamiliarity with tools such as Tableau for building dashboards + Ability to travel internationally without restriction + Demonstrate an unwavering passion for engineering products with exceptional reliability **Preferred Qualifications** + Master’s degree or PhD in engineering (mechanical, electrical, materials, etc.) or science (new college graduate) + Statistical experience such as Weibull, JMP, or familiarity with accelerated test models + Excel at building and maintaining reliability data dashboards on Tableau that drive informed decision making + Proficient in using Python, R or other languages to for extracting data and performing reliability analysis + Experience in adopting emerging AI/ML tools to streamline workflow and identify opportunities for automation + Excellent written and verbal communication skills for audiences ranging from technicians to senior executives + Familiarity with Failure Analysis techniques (Optical Microscopy, X-ray/CT, Scanning Electron Microscopy/Energy Dispersive Spectroscopy, etc.), and the ability to use failure analysis methodology to derive a root cause of failure + Experience in spirited collaboration between multidisciplinary teams of engineers and scientists to solve complex electromechanical / mechanical design challenges **Minimum Qualifications** + Bachelor’s degree in engineering (mechanical, electrical, materials, etc.) or science with 3+ of industry experience + Collaborative attitude and ability to cross-functionally work effectively + Familiarity with failure analysis methodology to identify root cause + Ability to make clear and concise slides/ presentations through Keynote, Excel, etc.
Senior Site Reliability Engineer - ASE / iCloud Apple IncSenior Site Reliability Engineer - ASE / iCloudCupertino, CAContribute to capacity planning, scale testing, and disaster recovery exercises Approach operational problems with a software engineering mindset.5+ years in a Infrastructure Ops, Site Reliability Engineering, or DevOps focused role. ASE Products Site Reliability teams are responsible for the reliability and performance of the server software stack that powers products like iCloud Photos, Mail, Drive, Backup and many more.
Senior Site Reliability Engineer, Apple Data Platform SRE / Apple Services Engineering Apple IncSenior Site Reliability Engineer, Apple Data Platform SRE / Apple Services EngineeringCupertino, CAThe Apple Data Platform (ADP) SRE Technical Lead partners with multiple SRE and engineering teams across the data platform - including teams responsible for Hadoop and HBase infrastructure, Spark, S3-compatible storage, and Airflow-orchestrated pipelines. Rather than owning a single vertical, this role sets the technical direction for how reliability is practiced across ADP: defining SLOs, establishing architectural review processes, developing shared tooling and automation, and ensuring that SRE principles are applied consistently as the platform scales.