Senior Site Reliability Engineer - Undersea Dominance ANDURIL INDUSTRIESSenior Site Reliability Engineer - Undersea DominanceCosta Mesa, CA$166,000–$250,000 / yearIf you're driven to build systems that last, thrive on deep technical challenges, and want to see your work directly shape how we design, build, and sustain complex platforms, you'll be helping build the future of digital shipbuilding and the next generation of maritime vehicles. Monitoring and Logging: Establish comprehensive monitoring and logging solutions using tools like ELK Stack (Elasticsearch, Logstash, Kibana), Prometheus, and Grafana to ensure the smooth operation of deployment environments.
Site Reliability Engineer (SRE) Pyramid, IncSite Reliability Engineer (SRE)Chicago, IL$50–$53 / hourFull timeBy applying to our jobs, you agree to receive calls, AI-generated calls, text messages, or emails from Pyramid Consulting, Inc. and its affiliates, and contracted partners. Provide a proactive approach to our clients’ workloads, anticipating failures, automating tasks, ensuring availability, and providing a great customer experience.
Service Reliability Engineer Pyramid, IncService Reliability EngineerReston, VAFull timeKey Responsibilities and Requirements: BA/BS in Computer Science, Computer Engineering or related technical discipline, or in place of 4-year degree, an equivalent industry internship or industry software engineering experience. Experience on public and/or private cloud resource provisioning and management through automation (Eg: ansible/terraform/puppet etc.).
Maintenance Reliability Engineer GMMaintenance Reliability EngineerBay City, MichiganTHIS INCLUDES DIRECT COMPANY SPONSORSHIP, ENTRY OF GM AS THE IMMIGRATION EMPLOYER OF RECORD ON A GOVERNMENT FORM, AND ANY WORK AUTHORIZATION REQUIRING A WRITTEN SUBMISSION OR OTHER IMMIGRATION SUPPORT FROM THE COMPANY (e.g., H-1B, OPT, STEM OPT, CPT, TN, J-1, etc.) . This includes direct company sponsorship, entry of GM as the immigration employer of record on a government form, and any work authorization requiring a written submission or other immigration support from the company (e.g., H1-B, OPT, STEM OPT, CPT, TN, J-1, etc).
Fielded Site Reliability Engineer ANDURIL INDUSTRIESFielded Site Reliability EngineerWaltham, MA$166,000–$220,000 / yearYou'll be first and second line of response for issues coming through our support channels and Anduril's customer-support pipeline (Tier 0 to Mission Success / Product Operations to SRE), acting as the deep-expertise backstop the rest of the funnel escalates to. To ensure your safety and help you navigate your job search with confidence, please keep the following critical points in mind: No Financial Requests: Anduril will never solicit payment or demand personal financial details (such as banking information, credit card numbers, or social security numbers) at any stage of our hiring process.
Site Reliability Engineer ANDURIL INDUSTRIESSite Reliability EngineerWaltham, MA$166,000–$220,000 / yearYou'll be first and second line of response for issues coming through our support channels and Anduril's customer-support pipeline (Tier 0 to Mission Success / Product Operations to SRE), acting as the deep-expertise backstop the rest of the funnel escalates to. To ensure your safety and help you navigate your job search with confidence, please keep the following critical points in mind: No Financial Requests: Anduril will never solicit payment or demand personal financial details (such as banking information, credit card numbers, or social security numbers) at any stage of our hiring process.
Site Reliability Engineer BasetenSite Reliability EngineerSan Francisco, CaliforniaYou'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable companies operating at the frontier of AI to bring cutting-edge models into production.
Senior Network Reliability Engineer, Incident Management Skylo TechnologiesSenior Network Reliability Engineer, Incident ManagementAt Senior NRE level you independently command bridge calls for Sev 1-4 incidents, support subject matter experts as incident coordinator and function as the communication lead during Sev 1 events, drive the execution of runbooks without supervision, manage the full incident lifecycle end-to-end in the ticketing system, and continuously improve the procedures you operate against. Work in close collaboration across multiple Skylo functions during incidents: RAN NRE, Core NRE, Cloud Infrastructure NRE, BOSS (BSS & OSS), Network Implementation, Product Engineering and Market teams, coordinating without creating confusion by maintaining a single source of truth on the bridge.
Site Reliability Engineer Rapinno TechSite Reliability EngineerPiscataway, New JerseyProficiency in high level script languages (Python preferred) as well as script environments like bash Experience with DevOps workflow automation (Jenkins, Ansible, Puppet). Strong experience with Installing IBM WebSphere MQ and creating multi instance Queue manager in AWS by using EBS/EFS volumes, creating MQ objects, clusters, channels etc.
Site Reliability Engineer with Java Knack SolutionsSite Reliability Engineer with JavaIrving, TexasRequired skills: AppDynamics, Splunk, Kibana/ELK, Kubernetes, Jenkins etc (2-3 years experience). Preferred Skills: APIGEE, Java, Cassandra, React.
Site Reliability Engineer (GKE) - Draper, UT ProdataKeySite Reliability Engineer (GKE) - Draper, UTCross-Functional Collaboration: Partner closely with backend developers and hardware engineering teams to define Service Level Indicators (SLIs), Service Level Objectives (SLOs), and manage error budgets. By leveraging automation, designing resilient distributed systems, and championing observability, you will ensure that our customers can secure and manage their access control systems without friction or failure.
Manufacturing Reliability Engineer (Austin, TX; 50% Travel To Milwaukee, WI) Diligent RoboticsManufacturing Reliability Engineer (Austin, TX; 50% Travel To Milwaukee, WI)Austin, TXWe're hiring a Manufacturing Reliability Engineer to own production test for our robots at our contract manufacturer: you'll design and run robust end-to-end test protocols, provision fleets of robots for production, and own the KPIs that define production quality. Broad technical knowledge across systems integration (ES integration), sensors (e.g., cameras, LiDAR, IMU or similar), embedded compute platforms, and networking/provisioning.
Senior Site Reliability Engineer II Juniper SquareSenior Site Reliability Engineer IIRemoteWe’re the Operations Partner trusted by 2,300+ GPs, unifying technology, data, and fund administration services into a single platform that helps GPs move faster, make better decisions, and scale with precision. Founder-led since 2014, backed by $350M+ in funding, and now 1,000+ employees strong, we’re building a company designed to shape the future of private markets for decades to come.
Customer Reliability Engineer - Airflow AstronomerCustomer Reliability Engineer - AirflowSan Francisco, CaliforniaRemote$125,000–$130,000 / yearBecause our customers are sophisticated organizations who need and expect high levels of expertise to help them keep mission critical uses of Apache Airflow working consistently, we look a little different from most support teams. Spend up to 20% of your time on side projects that contribute to Astronomer’s overall success, such as contributing to the open-source Airflow repository or developing Astronomer’s internal monitoring and alerting systems built on Airflow.
Lead Reliability Engineer - Mechanical Aligned Data CentersLead Reliability Engineer - MechanicalArizonaExperience: 10+ years of advanced mechanical industry experience (or advanced degree in Mechanical Engineering and 6+ years), specifically focused on hyperscale data center infrastructure, complex hydronic loops, and massive-scale chiller plants. Role Summary: Unlike traditional operational support roles, the Lead Reliability Engineer is the supreme technical authority for hydronic and mechanical infrastructure across Aligned's fleet.
Infrastructure Reliability Engineer STACK InfrastructureInfrastructure Reliability EngineerManassas, VirginiaResponsibilities include but are not limited to: Lead deep-dive investigations and RCAs for electrical infrastructure failures, including UPS systems, switchgear, breakers, relays, generators, grounding systems, STS behavior, VFD interactions, controls, and power quality disturbances. Partner with Operations, Engineering and Construction to review electrical design assumptions, protective schemes, equipment compatibility, and commissioning practices to identify long-term reliability risks prior to or following operational events.
Site Reliability Engineer ICONMA, LLCSite Reliability EngineerRemote, ALRemote$64.96–$69.96 / hourCollaboration: Work closely with developers, scientists, and infrastructure teams to deliver reliable platform services and translate operational needs into sustainable engineering solutions. GitOps and Deployment Automation: Support and improve GitOps workflows using ArgoCD to manage cluster and application configuration in a consistent, auditable manner.
Senior Site Reliability Engineer MicrosoftSenior Site Reliability EngineerRedmond, WA$119,800–$234,700 / yearJoining the Azure Cosmos DB team is a fantastic opportunity to work with incredibly talented engineers operating like a startup and be at the forefront of building and shaping the Livesite Automation and AI Ops stack in Cosmos DB and lead the path for broader adoption across Microsoft Azure. Microsoft’s Azure Data engineering team is leading the transformation of analytics in the world of data with products like databases, data integration, big data analytics, messaging & real-time analytics, and business intelligence.
NewSenior Reliability Engineer Amphenol HSIOSenior Reliability EngineerYocumtown, PAAmphenol Corporation is one of the world’s largest designers and manufacturers of electrical, electronic and fiber optic connectors and interconnect systems, antennas, sensors and sensor-based products and coaxial and high-speed specialty cable. Amphenol Corporation is one of the world's largest designers and manufacturers of electrical, electronic and fiber optic connectors and interconnect systems, antennas, sensors and sensor-based products and coaxial and high-speed specialty cable.
Production Support Analyst/SRE Reliability engineer Veterans Sourcing GroupProduction Support Analyst/SRE Reliability engineerRiverton, UTThis role supports Production Support / SRE (Reliability Engineering)functions—focused on system stability, incident resolution, automation, and improving platform reliability in a large-scale Linux environment. Troubleshoot incidents, perform root cause analysis, and resolve live issues.