Sr. Manager, Site Reliability Engineer

iCIMS Talent Acquisition
  • Holmdel, New Jersey
  • $150,000–$170,000 Per Year
  • Full-time
2 days ago

Job Description

Overview:

We are seeking a Sr Manager of Site Reliability Engineering (SRE) to lead and continue developing our global SRE organization across the US, Ireland, India, and strategic engineering partners. 

 

This role is responsible for advancing a modern SRE operating model that combines centralized reliability capabilities with product-aligned SRE support. The Sr Manager will partner closely with Product Engineering, Cloud Engineering, Operations, Security, DBA, and other technical teams to improve reliability, strengthen operational practices, and help development teams implement consistent engineering standards. 

 

This is a highly technical leadership role requiring strong experience across Site Reliability Engineering, AWS cloud architecture, observability, cloud engineering, automation, incident and problem management, and FinOps. The successful candidate will be able to evaluate technical decisions across reliability, scalability, security, operational complexity, and cloud financial impact. 

About Us:

ICIMS is a leading enterprise hiring platform that combines the scale and reliability of enterprise software with the transformative power of AI. Thousands of organizations across more than 200 countries and territories trust ICIMS to find and hire the people who shape their future and drive their business forward. Powered by insights from billions of hiring interactions, continuous AI innovation, and a highly extensible platform, ICIMS helps organizations turn talent acquisition into a competitive advantage. For more than 25 years, ICIMS has delivered end-to-end hiring solutions that improve recruiting efficiency, reduce costs and create exceptional candidate experiences. 

 

ICIMS helps solve one of the biggest challenges businesses face today: building a workforce that can adapt, scale, and perform in an increasingly competitive and unpredictable talent market. We uniquely do that by combining enterprise-grade hiring technology, AI-powered insights and automation, and connected talent experiences to help organizations improve hiring outcomes while driving measurable impact. 

Responsibilities:

Leadership & Strategy 

  • Lead and develop a globally distributed SRE organization across multiple regions and time zones. 
  • Define and execute the SRE strategy, operating model, priorities, and technical direction. 
  • Establish a hybrid SRE model combining centralized Reliability Enablement capabilities with product-aligned SRE support. 
  • Define clear responsibilities and decision rights across SRE, Product Engineering, Operations, Cloud Engineering, DBA, and other technical functions. 
  • Build strong technical leadership, product ownership, regional handoffs, and knowledge-sharing practices across the global organization. 
  • Develop engineers and technical leaders through coaching, mentorship, and clear technical and career expectations. 
  • Drive a culture focused on proactive reliability engineering rather than reactive operational support. 

 

Reliability Engineering & Product Alignment 

  • Partner with Product Engineering teams to understand product architecture, service dependencies, reliability risks, and operational requirements. 
  • Establish and mature SLIs, SLOs, error budgets, service-health measures, and operational-readiness standards where appropriate. 
  • Identify systemic reliability issues before they become customer-impacting incidents. 
  • Translate production learnings, recurring failures, and RCAs into prioritized engineering improvements. 
  • Establish reusable reliability patterns, tooling, automation, runbooks, and engineering practices that can be adopted across product teams. 
  • Ensure SRE supports product teams without replacing Product Engineering ownership of application functionality and defects. 

 

Problem Management & Operational Excellence 

  • Provide technical leadership during significant and complex production incidents. 
  • Partner with Operations and engineering teams to strengthen incident response, escalation, restoration, and recovery practices. 
  • Lead the continued development of structured problem management and root-cause analysis practices. 
  • Connect recurring incidents and operational risks to visible, prioritized corrective actions. 
  • Improve operational readiness, capacity planning, performance management, and service resiliency. 
  • Reduce repetitive operational work and manual intervention through automation and engineering. 

 

Observability & Reliability Enablement 

  • Lead the development and adoption of common observability standards across logging, metrics, tracing, dashboards, monitoring, and alerting. 
  • Drive enterprise observability strategy and governance across platforms including Grafana, OpenTelemetry, Sumo Logic, New Relic, CloudWatch, and related technologies. 
  • Establish practical service blueprints and reusable observability patterns for development teams. 
  • Improve alert quality, service visibility, dependency awareness, and actionable monitoring. 
  • Ensure observability capabilities support both real-time incident response and longer-term reliability improvement. 

 

Operational FinOps & Cloud Financial Management 

  • Incorporate cloud financial awareness into architecture, reliability, and engineering decisions. 
  • Partner with Cloud Engineering, Finance, Product, and engineering leadership to improve cloud cost ownership and accountability. 
  • Identify opportunities for rightsizing, workload optimization, storage efficiency, and removal of unnecessary cloud consumption. 
  • Understand AWS pricing and commitment constructs including Savings Plans, Reserved Instances, licensing considerations, and consumption-based services. 
  • Support effective cost allocation, tagging, forecasting, reporting, and cloud financial governance. 
  • Evaluate technical decisions using both engineering and financial considerations while ensuring cost optimization does not introduce unacceptable reliability or performance risk. 

 

AWS & Cloud Engineering  

  • Provide senior technical leadership for complex cloud environments, with AWS as the primary platform. 
  • Partner with Cloud Engineering on architecture, resiliency, automation, networking, security, governance, and platform standards. 
  • Review and guide architectures involving AWS technologies such as ECS, ECR, EC2, RDS, Aurora, S3, DynamoDB, OpenSearch, SQS, SNS, Kinesis, IAM, AWS Organizations, and cloud networking. 
  • Evaluate architecture across availability, scalability, performance, security, recoverability, operational complexity, and cost. 
  • Drive Infrastructure as Code, automated provisioning, standardized cloud patterns, and policy-based governance
Qualifications:
  • 10+ years of experience across Site Reliability Engineering, Cloud Engineering, Platform Engineering, DevOps, Infrastructure Engineering, or related technical disciplines. 
  • 5+ years of technical or engineering leadership experience, including responsibility for technical strategy, team leadership, and organizational outcomes. 
  • Experience leading and collaborating with distributed technical teams across multiple regions and time zones. 
  • Strong technical knowledge of AWS and experience supporting or designing complex enterprise cloud environments. 
  • Broad understanding of cloud computing, containers, networking, storage, databases, security, identity, monitoring, and governance. 
  • Strong understanding of highly available, scalable, resilient, and distributed production systems. 
  • Experience with Infrastructure as Code and automated infrastructure delivery, preferably Terraform. 
  • Strong understanding of modern observability practices including logs, metrics, traces, monitoring, dashboards, and alerting. 
  • Demonstrated experience with incident response, technical escalation, root-cause analysis, and problem management. 
  • Strong understanding of FinOps and cloud financial management principles, including optimization, allocation, forecasting, and cost accountability. 
  • Ability to evaluate architecture from both technical and financial perspectives. 
  • Strong communication and collaboration skills with the ability to influence engineers, architects, Product leaders, Finance, Security, and executive stakeholders. 
EEO Statement:

iCIMS is a place where everyone belongs. We celebrate diversity and are committed to creating an inclusive environment for all employees. Our approach helps us to build a winning team that represents a variety of backgrounds, perspectives, and abilities. So, regardless of how your diversity expresses itself, you can find a home here at iCIMS.

 

We are proud to be an equal opportunity and affirmative action employer. We prohibit discrimination and harassment of any kind based on race, color, religion, national origin, sex (including pregnancy), sexual orientation, gender identity, gender expression, age, veteran status, genetic information, disability, or other applicable legally protected characteristics. If you would like to request an accommodation due to a disability, please contact us at careers@icims.com.

Compensation and Benefits:

We accept applications for this position on an ongoing basis until the position is filled. Applications will be reviewed as they are received, and qualified candidates may be contacted throughout the posting period. 

 

The anticipated base salary range for this position is $150,000 – $170,000. In addition, the estimated on-target earnings (“OTE”), which includes base salary and commissions, is $170,000-$200,000.

  

Actual compensation will depend on various job-related factors, including but not limited to, location, experience, and job qualifications. This range aligns with our commitment to equitable and transparent compensation practices, as required by applicable law. 

 

Competitive health and wellness benefits include medical, dental, vision, 401(k), dependent care, short term and long-term disability, life and AD&D insurance, bonding and parental leave, mindfulness resources, an open vacation policy, sick days, paid holidays, quiet hours each workday, and tuition reimbursement. Benefits and eligibility may vary by location, role, and tenure.  Learn more here: https://careers.icims.com/benefits

Numbers & Facts

LocationHolmdel, New Jersey
Job TypeFull-time
Salary$150,000–$170,000 Per Year

Skills

  • Amazon Elastic Compute Cloud (EC2)unmatched
  • Amazon Simple Notification Service (SNS)unmatched
  • Amazon Simple Storage Service (S3)unmatched
  • Amazon Web Services (AWS)unmatched
  • Artificial Intelligence (AI)unmatched
  • Automationunmatched
  • Blueprintsunmatched
  • Budgetingunmatched
  • Business Operationsunmatched
  • Capacity Managementunmatched
  • Capacity and Performance Managementunmatched
  • Cloud Architectureunmatched
  • Cloud Computingunmatched
  • Coachingunmatched
  • Communication Skillsunmatched
  • Compensation and Benefitsunmatched
  • Corrective Actionunmatched
  • Cost Allocationunmatched
  • Cost Controlunmatched
  • Cost Forecastingunmatched
  • Database Administrationunmatched
  • DevOpsunmatched
  • Distributed Computingunmatched
  • Diversityunmatched
  • Enterprise Applicationsunmatched
  • Establish Prioritiesunmatched
  • Financeunmatched
  • Financial Managementunmatched
  • Financial Reportingunmatched
  • Forecastingunmatched
  • High Availabilityunmatched
  • Identify Issuesunmatched
  • Incident Managementunmatched
  • Incident Responseunmatched
  • Information/Data Security (InfoSec)unmatched
  • Leadershipunmatched
  • Legalunmatched
  • Licensingunmatched
  • Machine Toolunmatched
  • Mentoringunmatched
  • Metricsunmatched
  • Multiplatform/Cross-Platformunmatched
  • Network Securityunmatched
  • Operational Improvementunmatched
  • Operational Measurementunmatched
  • Operational Supportunmatched
  • Operations Security (OPSEC)unmatched
  • Pricingunmatched
  • Product Engineeringunmatched
  • Product Supportunmatched
  • Production Systemsunmatched
  • Quality Managementunmatched
  • Reliability Engineeringunmatched
  • Reporting Dashboardsunmatched
  • Right-Sizingunmatched
  • Riskunmatched
  • Root Cause Analysisunmatched
  • SUMOunmatched
  • Simple Queue Service (SQS)unmatched
  • Software Engineeringunmatched
  • Team Buildingunmatched
  • Team Lead/Managerunmatched
  • Team Playerunmatched
  • Technical Analysisunmatched
  • Technical Leadershipunmatched
  • Technical Strategyunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder