Principal Site Reliability Engineer, Google Cloud SaviyntPrincipal Site Reliability Engineer, Google CloudMilpitas, CA$240,000–$250,000 / yearDesign and implement shared Event-Driven Architecture components and messaging platforms using technologies like Kafka or Google Pub/Sub that product teams can easily utilize. You will focus on creating reusable, reliable, and scalable solutions that abstract away complexity, enabling other teams to focus on their core business logic and deliver features faster in a multi-cloud environment.
Staff Site Reliability Engineer ANDURIL INDUSTRIESStaff Site Reliability EngineerCosta Mesa, CA$191,000–$253,000 / yearTo ensure your safety and help you navigate your job search with confidence, please keep the following critical points in mind: No Financial Requests: Anduril will never solicit payment or demand personal financial details (such as banking information, credit card numbers, or social security numbers) at any stage of our hiring process. Demonstrated experience designing and owning reliability infrastructure that multiple engineering teams depend on, including observability platforms, deployment systems, or incident management tooling at meaningful scale.
Senior Engineer, Reliability Engineering Analog DevicesSenior Engineer, Reliability EngineeringWilmington, MassachusettsThe successful candidate will work closely with reliability engineers, technicians, product teams, and senior hardware developers to deliver robust test solutions and support reliability qualification activities across a broad range of technologies and products. ADI combines analog, digital, AI, and software technologies into solutions that combat climate change, reliably connect humans and the world, and help drive advancements in automation and robotics, mobility, healthcare, energy and data centers.
Senior Network & Site Reliability Engineer AlembicSenior Network & Site Reliability EngineerSan Francisco, CaliforniaYou'll design and operate the global network and reliability layer behind one of the world's fastest private supercomputers — the fabric powering distributed compute, ML workloads, real-time analytics, and mission-critical enterprise systems. Extensive hands-on experience with network devices (firewalls, switches, load balancers) and large-scale architectures and protocols — BGP, QoS, MPLS, and IPsec VPNs.
NewManufacturing Industrial Engineer Ashley Furniture Industries, Llc.Manufacturing Industrial EngineerHolmen, WI$65,000–$75,000 / yearThis positionis located inan active industrial manufacturing and distribution center with moderate to high noise levels, temperature variations, and material handling equipment in operation. We will not pay any placement,referralor other fees to any search firms unless we have agreed otherwise in a valid, written agreement for the specific position posted and signed by an authorized representative of Ashley Furniture Industries.
Reliability Engineer Pinnacle Reliability IndiaReliability EngineerPasadena, TexasThe Impact of a Reliability Data AnalystAs a Reliability Data Analyst (RDA), you’ll be exposed to programs like Mechanical Integrity (MI), Risk Based Inspection (RBI), Reliability Centered Maintenance (RCM), Quantitative Reliability Optimization (QRO), technology development and implementation, and groundbreaking research and development opportunities. Ability to walk, stand, sit, kneel, push, stoop, reach above the shoulder, grasp, pull, bend repeatedly, climb stairs, identify colors, hear with aid, see, write, count, read, speak, analyze, lift and carry under 30 lbs., and perceive depth.
Senior Site Reliability Engineer -AI Infrastructure Operations NscaleSenior Site Reliability Engineer -AI Infrastructure OperationsHouston, Texas$170,000–$265,000 / yearStrong software engineering skills (Python, Go, or similar); you build tools other engineers adopt,• not scripts that run once and rot.• Deep command of Linux, networking, and distributed systems, plus the judgment to know where• the real failure modes hide.• Experience running AI or GPU workloads, or high-performance computing (HPC); if not, the depth to• get there fast.• Reliability practices you put in place that outlasted you: SLOs, observability and alerting at scale,• incident process, on-call that people can actually live with.•
Reliability Engineer IntelReliability EngineerUs, MassachusettsDefine and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) for compute, memory, storage, network, power, and cooling subsystems. Mission: Define and own the pod-level reliability specifications that ensure the availability, resilience, and serviceability of a large-scale data center across hardware, thermal, and operational dimensions.
Senior Site Reliability Engineer StordSenior Site Reliability EngineerStord’s end-to-end commerce solutions combine best-in-class omnichannel fulfillment and shipping with leading technology to ensure fast shipping, reliable delivery promises, easy access to more channels, and improved margins on every order. The SRE team is small, fast-moving, and owns the infrastructure that keeps that platform running, primarily on Google Cloud Platform, across GKE, Cloud Run, AlloyDB, and the networking that connects our services.
Failure Analysis and Reliability Engineer Advanced Micro Devices, IncFailure Analysis and Reliability EngineerSan Jose, CaliforniaThis role offers the opportunity to build deep hands-on failure analysis expertise, work with advanced products and analytical tools, and contribute directly to product quality, reliability, and customer success. This is a hands-on engineering position suited for someone who enjoys solving complex technical problems, working across teams, and driving issues from initial investigation through root-cause conclusion.
Site Reliability Engineer - AWS SitusAMCSite Reliability Engineer - AWSRemoteYou will maintain operational coverage of environments and continuously look for optimization, reengineering, and efficiency to ensure the products running on the cloud as a SaaS offering to our clients are stable and reliable. Leveraging various DevOps approaches which include but not limited to CI/CD processes, working closely with development teams to ensure a manageable and secure migration of change into the production environment.
Senior Site Reliability Engineer - Managed Kubernetes LambdaSenior Site Reliability Engineer - Managed KubernetesSan Francisco, CaliforniaOur investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove. *Note: This position requires presence in our San Francisco, San Jose, or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
Maintenance Reliability Engineer Howmet AerospaceMaintenance Reliability EngineerLa Porte, INCoordinate all phases of assigned projects from concept, scope development, written specifications, bid solicitations, related bid meetings, equipment and facilities engineering/design, manufacturing, and installation through final start-up, acceptance and project closure. Partner with the Human Resources function to establish, coordinate and monitor department-specific people systems (performance management, reward and recognition, compensation, employee engagement, employee development, succession planning, hiring process, new employee onboarding, etc).
Staff AI Platform & Reliability Engineer eJamStaff AI Platform & Reliability EngineerSanta Ana, CAFluency with agentic coding tools (Claude Code, Codex, Cursor or comparable), including the ability to scope work for them, review their output critically and maintain architectural coherence across AI-assisted changes. We are seeking a Staff Engineer to take ownership of the AI generation services, provider integrations, billing integrity, tenant security and overall platform reliability underpinning created.ai.
Site Reliability Engineer CVS HealthSite Reliability EngineerWoonsocket, Rhode IslandThe ideal candidate will leverage AIOps practices and AI-powered tools to improve observability, automate incident response, reduce operational noise, and enable data-driven decision making across distributed systems. We are seeking a highly skilled Software Development Engineer, Site Reliability Engineering (SRE), for Retail and Pharmacy platforms to drive reliability, scalability, and operational excellence.
Distribution Reliability Engineer Asplundh Engineering Services, LLCDistribution Reliability EngineerAtlanta, GAFull timeAbout Us Asplundh Engineering Services (AES) brings over 100 years of expertise to deliver top-notch engineering, design, planning, and testing services for transmission, distribution, substation, and renewable infrastructure markets across the U.S. An Asplundh Company As part of the Asplundh family, AES benefits from extensive resources and a century of experience. As a Mid‑Level Distribution Engineer with Asplundh Engineering Services (AES), you will support safe, reliable electrical infrastructure by designing upgrades, evaluating system performance, and partnering closely with utility clients across Georgia.
Site Reliability Engineer HireTalentSite Reliability EngineerChicago, ILYou ll work closely with peers and senior engineers to manage and evolve our infrastructure landscape, support cloud migrations, and independently deliver on key automation and integration projects. Develop and maintain file transfer processes across 100+ internal flows for critical daily financial file transfers.
Principal Engineer, Rotating Equipment CP2 LNG OperationsPrincipal Engineer, Rotating EquipmentAdditionally, this individual shall provide, as a minimum, intermediate level Reliability Engineering support to the Operations team with focus in the following areas: Root Cause Failure Analysis (RCFA): With expertise in relevant techniques to adequately lead RCFA exercises for failures relating to plant equipment and/or processes. Using reliable, proven technology in an innovative plant design configuration, Venture Global’s modular, mid-scale plant design will replace traditional designs as it allows for the same efficiency and operational reliability at significantly lower capital cost.
Senior Site Reliability Engineer (Aws Cloud) WizelineSenior Site Reliability Engineer (Aws Cloud)Romania, PAWe're looking for a Senior SRE Engineer with solid experience in engineering and architecture on AWS, and a strong ability to automate infrastructure and configuration in complex projects. Wizeline is a global AI-native technology solutions provider that develops cutting-edge, AI-powered digital products and platforms.
Lead Site Reliability Engineer, Vice President Morgan StanleyLead Site Reliability Engineer, Vice PresidentNew York, New York$150,000–$190,000 / yearOur values - putting clients first, doing the right thing, leading with exceptional ideas, committing to diversity and inclusion, and giving back - aren’t just beliefs, they guide the decisions we make every day to do what's best for our clients, communities and more than 80,000 employees in 1,200 offices across 42 countries. Test and tune network, hardware, and software configurations to maximize performance * Interface with different teams like IT Dev managers, Infrastructure teams and lead as a Subject Matter Expert (SME) for the application(s) supported.