Senior Site Reliability Engineer Castleton Commodities International LLCSenior Site Reliability EngineerStamford, CTThe Senior Site Reliability Engineer will also lead efforts to define and validate recovery objectives (RTO/RPO), design and implement Business Continuity / Disaster Recovery (BCP/DR) plans, and coordinate structured testing to ensure readiness. The Senior Site Reliability Engineer is responsible for improving the reliability, availability, scalability, and operational excellence of our critical infrastructure platforms and services.
Data Reliability Engineer Damco Solutions IncData Reliability EngineerBasking Ridge, NJResponsibilities: Proactively monitor PostgreSQL RDS instances for performance, availability, and resource utilization (CPU, memory, storage, connections) using established monitoring tools (e.g., CloudWatch, Prometheus). As a Data Reliability Engineer II, you will play a crucial role in developing, optimizing, and managing several large data lakes and data warehouses, comprising data from multiple disparate sources.
Site Reliability Engineer (SRE) Production Support-W2-ONSITE RAPS Consulting IncSite Reliability Engineer (SRE) Production Support-W2-ONSITEIselin, NJDemonstrated client facing experience as a core responsibility, including direct communication with customers, requirement gathering, status updates, and expectation management — beyond internal coordination with development teams . The ideal candidate will be responsible for monitoring production environments, troubleshooting incidents, ensuring system availability, and supporting reliability and automation initiatives.
Site Reliability Engineer KalshiSite Reliability EngineerNew York City, New YorkKalshi fought for years and legalized prediction markets in the US for the first time in history, is currently the fastest growing financial market in America, and has thousands of markets across politics, economics, financials, weather, tech, AI, culture and more. As a member of Kalshi’s engineering team, you’ll help build the next-generation financial ecosystem—similar in ambition to creating an NYSE or CME from scratch.
Site Reliability Engineer PicoSite Reliability EngineerNew York City, New YorkPico Site Reliability Engineering is a customer-faced group engaged with our customers, development and production teams to ensure customer success with implementing Redline products. The role holder will be expected to work whatever hours are necessary for the performance of this role (recognizing that it involves multiple jurisdictions/geographies including but not limited to EMEA, USA and APAC).
NewReliability Engineer OFIReliability EngineerBayonne, NJ$99,000–$132,000 / yearPosition Summary Reporting to the Plant Manager, The Reliability Engineer will lead and collaborate with cross-functional, multidisciplinary teams ensure systems, equipment, and processes operate consistently, efficiently and safely by identifying risks, analyzing failures, and implementing preventive strategies. In addition, the ideal candidate will have effectively demonstrated an ability to continuously deliver cost savings opportunities by efficiency and downtime reduction, build strategic approach and methodologies, establish better services, and demonstrate leadership.
Site Reliability Engineer Rapinno TechSite Reliability EngineerPiscataway, New JerseyProficiency in high level script languages (Python preferred) as well as script environments like bash Experience with DevOps workflow automation (Jenkins, Ansible, Puppet). Strong experience with Installing IBM WebSphere MQ and creating multi instance Queue manager in AWS by using EBS/EFS volumes, creating MQ objects, clusters, channels etc.
Staff Reliability Engineer ZT SystemsStaff Reliability EngineerSecaucus, NJAct as the internal consultant on all reliability matters and interface with program management, vendors, and design engineering (as necessary) on key reliability programs/issues; supporting the Software/script development needs of the reliability team. Please be aware that certain positions may require the applicant to either 1) be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or 2) be eligible to obtain an export control license or license exception from the Bureau of Industry and Science & U.S. Department of Commerce.
Senior DevOps Engineer/Site Reliability Engineer-East Coast Stellar CyberSenior DevOps Engineer/Site Reliability Engineer-East CoastNew York, New YorkRemoteThis role combines DevOps engineering practices with SRE principles to improve scalability, resiliency, operational efficiency, and platform performance across production environments. This is a hands-on senior-level role focused on building, operating, and scaling reliable cloud-native infrastructure and distributed data platforms.
NewSenior Site Reliability Engineer Cross RiverSenior Site Reliability EngineerFort Lee, NJRemote$160,000–$200,000 / yearOur technology and capital solutions power payments, cards, lending, and digital asset capabilities that move money safely, instantly, and inclusively — trusted by leading fintechs, enterprises, and disruptors across the globe. The ideal candidate brings deep expertise in building and maintaining production infrastructure, establishing DevOps best practices, and driving operational excellence across engineering teams.
NewReliability Engineer 3 (Observability Specialist) U.S. BancorpReliability Engineer 3 (Observability Specialist)New York, NY$98,175–$115,500 / yearThe engineer provides senior technical guidance, identifies observability gaps through incident and performance analysis, drives continuous improvement, and helps ensure teams have the data, processes, and operating discipline needed to detect issues earlier, reduce customer impact, and improve overall service reliability. Lead the definition, documentation, and ongoing refinement of critical user journeys in partnership with product owners, engineering teams, SRE, operations, and business stakeholders to ensure observability practices are aligned to customer experience, business outcomes, and operational risk.
Site Reliability Engineer, Observability RippleSite Reliability Engineer, ObservabilityNew York, NY$160,000–$200,000 / yearWith more than 40 years of experience supporting some of the world's largest and most sophisticated companies, Ripple Treasury integrates a treasury command center into Ripple's technology stack-giving corporates the ability to move, manage, and optimize liquidity in real-time, across traditional and digital assets, under one expanded umbrella. You will spend the majority of your time doing hands-on observability and reliability engineering work: building instrumentation, designing alert configurations, authoring Terraform, and troubleshooting production systems.
Senior Site Reliability Engineer, Observability RippleSenior Site Reliability Engineer, ObservabilityNew York, NY$160,000–$200,000 / yearWith more than 40 years of experience supporting some of the world's largest and most sophisticated companies, Ripple Treasury integrates a treasury command center into Ripple's technology stack-giving corporates the ability to move, manage, and optimize liquidity in real-time, across traditional and digital assets, under one expanded umbrella. You will spend the majority of your time doing hands-on observability and reliability engineering work: building instrumentation, designing alert configurations, authoring Terraform, and troubleshooting production systems.
NewSite Reliability Engineer, RL Infra Thinking Machines Lab IncSite Reliability Engineer, RL InfraNew York, NY$350,000–$475,000 / yearYoull work alongside the engineers building the RL training and rollout systems, and with the research teams running experiments on top of them, to make every layer of the stack more robust and resilient: from the rollout and inference-serving layer, through weight synchronization between trainers and samplers, to the schedulers deciding how GPUs are shared across simultaneous jobs. Design and implement monitoring and observability across the full RL loop - rollout generation, environment and tool execution, reward computation, and weight synchronization between trainers and samplers - so failures and slowdowns are caught close to their source.
NewSite Reliability Engineer, Model Post Training Thinking Machines Lab IncSite Reliability Engineer, Model Post TrainingNew York, NY$350,000–$475,000 / yearYoull be trusted to define what reliable means for post-training workloads, push back on launches that put stability or correctness at risk, and lead incident response for issues that span multiple teams and layers of the stack. Design and implement monitoring and observability across the full training path - data loading, forward/backward passes, optimizer steps, checkpointing, and sampling - with enough signal to distinguish an infrastructure failure from a training or modeling issue.
Staff Site Reliability Engineer DiligentStaff Site Reliability EngineerNew York, NY$131,000–$164,000 / yearPartner with network and security teams to manage firewalls, VPNs, storage, and load balancers (F5 BIG‑IP, AVI/NSX Advanced Load Balancer) for highly available services. The Diligent One Platform gives practitioners, the C-Suite and the board a consolidated view of their entire GRC practice so they can more effectively manage risk, build greater resilience and make better decisions, faster.
Site Reliability Engineer, Pragma MarketAxess Holdings, Inc.Site Reliability Engineer, PragmaNew York, NY$175,000–$230,000 / yearYou'll help drive modernization efforts, strengthen SRE practices, partner with internal trading desks and engineering teams, and work directly with clients and external partners. We are seeking a Site Reliability Engineer to join a team responsible for the production and non-production environments supporting multiple in-house trading systems.
Senior Site Reliability Engineer DatavantSenior Site Reliability EngineerNew York, NY$168,000–$200,000 / yearThe estimated total cash compensation range for this role is: $168,000—$200,000 USD To ensure the safety of patients and staff, many of our clients require post-offer health screenings and proof and/or completion of various vaccinations such as the flu shot, Tdap, COVID-19, etc. Guided by our mission to make the world's health data secure, accessible and actionable, we provide critical data solutions for organizations across the healthcare ecosystem - including providers, health plans, researchers, and life sciences companies.
Lead Site Reliability Engineer KontaktLead Site Reliability EngineerNew York, NY$200,000–$250,000 / yearBacked by Goldman Sachs and trusted by leading health systems including HCA Healthcare, Sutter Health, AdventHealth, Trinity Health, Northwell Health, Cleveland Clinic, and the U.S. Department of Veterans Affairs, we're pioneering the next generation of healthcare operations. You will lead and scale the SRE team to ensure our infrastructure stays ahead of demand, operates efficiently, and meets the needs of our growing healthcare customers.
Senior Lead Platform Reliability Engineer Wells Fargo & CoSenior Lead Platform Reliability EngineerIselin, NJ$159,000–$305,000 / yearThis role is designed for highly experienced infrastructure engineers who possess deep technical expertise in one core platform discipline (Network, Middleware, Database, or Storage) and have demonstrated experience collaborating across at least one additional infrastructure domains (Enterprise Tools, Cloud, Observability Tools). They are accountable for execution of all applicable risk programs (Credit, Market, Financial Crimes, Operational, Regulatory Compliance), which includes effectively following and adhering to applicable Wells Fargo policies and procedures, appropriately fulfilling risk and compliance obligations, timely and effective escalation and remediation of issues, and making sound risk decisions.