ISO New England Inc. logo

Senior Site Reliability Engineer

ISO New England Inc.
  • Holyoke, MA
  • $134,000–$170,000 Per Year
6 days ago

Job Description

The Senior Site Reliability Engineer (SRE) is a hands-on engineering role responsible for improving the reliability, observability, performance, and operational efficiency of ISO New England's IT services. The SRE works across infrastructure, platform, cyber security, and application teams to reduce operational toil, improve service resilience, and implement scalable automation solutions.

This role has a strong emphasis on observability engineering, automation, Splunk administration, and Infrastructure as Code (IaC). The ideal candidate will possess hands-on experience with Splunk or demonstrate a strong willingness to develop expertise in the platform. Experience with Terraform, automation technologies such as Python and PowerShell, and the ability to leverage AI-assisted development tools to accelerate engineering solutions are key components of the role.

What we offer you:

A stable, mission-driven workplace where your impact truly mattersA highly engaged work environment that values inclusion, collaboration, and employee safety and wellbeingCompetitive compensation with a base salary + performance bonusRobust benefits package, including:

Enhanced 401(k) and financial planning supportTuition reimbursement and professional developmentWellness programs, including an onsite gymFlexible work hoursEmployee Business Networks

Free coffee at our onsite café

Hybrid work environment (3 days/week onsite)Distance-based relocation assistance available

How you will make an Impact

Build and maintain observability, monitoring, logging, alerting, and telemetry platforms (e.g., Splunk, Dynatrace, PRTG, OpsGenie, StatusPage)Administer, maintain, automate, and continuously improve the Splunk platform, including data onboarding, indexing, search performance, dashboards, access controls, health monitoring, platform scalability, and operational workflowsDevelop and automate Splunk onboarding, configuration, monitoring, and operational workflows to improve platform reliability and reduce administrative overheadDevelop meaningful KPIs and dashboards for business and IT service healthEngineer and implement resilience patterns including HA, DR, and automated failoverPartner with infrastructure and application teams to plan and execute resilience testing and failover exercises to validate recovery capabilities and observability coverageConduct performance testing, capacity modeling, forecasting, and right-sizingParticipate in major incident response activities, providing technical expertise to accelerate service restoration and identify reliability improvementsIdentify, prioritize, and eliminate manual operational toil through automation, targeting workflows, runbooks, alerting, platform administration, service management processes, and KPI collection, with a bias toward scalable and repeatable engineering solutionsDesign, develop, maintain, and support automation solutions, integrations, and operational tooling using Python, PowerShell, Bash, or similar technologies to improve reliability, reduce manual effort, and enhance operational efficiencyDesign, deploy, and manage infrastructure using Terraform and Infrastructure as Code (IaC) practices, including observability platforms, infrastructure services, and supporting technology stacks, with a focus on consistency, repeatability, and operational sustainabilityIdentify gaps in observability coverage and drive engineering solutions to close themCollaborate with architecture and application teams to ensure production readinessLeverage AI-assisted development tools to accelerate automation initiatives while reviewing, validating, troubleshooting, and refining generated code to ensure reliability, security, maintainability, and operational effectivenessReduce repeat incidents by engineering permanent fixes and driving continuous improvement

What we are looking for

5+ years of experience in SRE, DevOps, systems engineering, platform engineering, or IT operationsExperience with enterprise monitoring and observability platforms. Hands-on experience with Splunk is strongly preferred. Candidates without direct Splunk experience must demonstrate a strong willingness and aptitude to develop expertise in Splunk administration, engineering, and automation.Experience designing, deploying, or managing infrastructure using Terraform and Infrastructure as Code (IaC) practicesStrong scripting and automation experience using Python, PowerShell, Bash, or similar technologies, including the development of operational tooling, integrations, and workflow automation in production environmentsDemonstrated experience designing, developing, and supporting automation solutions that measurably reduced manual operational effort in an enterprise environmentAbility to read, understand, review, troubleshoot, and refine code produced by engineering teams or AI-assisted development platformsKnowledge of distributed systems, networking, enterprise infrastructure, and cloud platformsFamiliarity with SRE principles including SLOs, error budgets, observability, and toil reductionAbility to analyze and troubleshoot complex technical systemsPreferred QualificationsExperience in mission-critical, highly available, or regulated environmentsExperience utilizing AI-assisted development tools to accelerate automation, operational engineering, or platform management activitiesKnowledge of ITIL processes and/or SRE best practicesExperience with performance testing, capacity planning, resilience testing, or disaster recovery validation

This employer will not sponsor applicants for work visas for this position (ex: H-1B, F-1/CPT/OPT, O-1, E-3, TN, J, etc.).

The expected salary range for this position is $134,000 - $170,000 per year, for a Senior to Lead level candidate. This role is also eligible for an annual performance bonus, comprehensive health insurance (medical, dental and vision), flexible spending and health savings accounts, a 401(k) plan with generous employer contributions and a student debt benefit, life and AD&D insurance, disability insurance, critical illness and hospital indemnity benefits, paid time off, paid leave, a wellness program, an employee assistance program and other great company perks.#LI-HYBRID

Numbers & Facts

LocationHolyoke, MA
IndustryEnergy and Utilities
Salary$134,000–$170,000 Per Year
Company Size1,000 to 1,499 employees
Year Founded1996
Websitehttps://www.iso-ne.com/

About Company

ISO New England helps protect the health of New England's economy and the well-being of its people by ensuring the constant availability of electricity, today and for future generations. ISO New England meets this obligation in three ways: by ensuring the day-to-day reliable operation of New England's bulk power generation and transmission system, by overseeing and ensuring the fair administration of the region's wholesale electricity markets, and by managing comprehensive, regional planning processes.

Skills

  • Access Controlunmatched
  • Administrative Managementunmatched
  • Analysis Skillsunmatched
  • Applications Securityunmatched
  • Artificial Intelligence (AI)unmatched
  • Automationunmatched
  • Bash Scriptingunmatched
  • Budgetingunmatched
  • Business Servicesunmatched
  • Capacity Managementunmatched
  • Cloud Computingunmatched
  • Continuous Improvementunmatched
  • DevOpsunmatched
  • Disaster Recoveryunmatched
  • Distributed Computingunmatched
  • Establish Prioritiesunmatched
  • Failoverunmatched
  • Financial Planningunmatched
  • Forecastingunmatched
  • High Availabilityunmatched
  • ISO (International Organization for Standardization)unmatched
  • ITIL (IT Infrastructure Library)unmatched
  • Identify Issuesunmatched
  • Incident Responseunmatched
  • Information Technology & Information Systemsunmatched
  • Internet Securityunmatched
  • Machine Toolunmatched
  • Multiplatform/Cross-Platformunmatched
  • Onboardingunmatched
  • Operational Improvementunmatched
  • Operational Measurementunmatched
  • Operations Planningunmatched
  • Performance Metricsunmatched
  • Performance Testingunmatched
  • Process Improvementunmatched
  • Process Managementunmatched
  • Productivity Modelunmatched
  • Programming Toolsunmatched
  • Python Programming/Scripting Languageunmatched
  • Reimbursementunmatched
  • Reliability Engineeringunmatched
  • Reporting Dashboardsunmatched
  • Scripting (Scripting Languages)unmatched
  • Splunkunmatched
  • Systems Administration/Managementunmatched
  • Systems Engineeringunmatched
  • Technical Supportunmatched
  • Telemetryunmatched
  • Testingunmatched
  • Windows PowerShellunmatched

Be found by employers

5,500+ employers search our resume database daily. Add yours to get found by recruiters looking for candidates like you.

Level up your application

Professional resume templates

Browse dozens of recruiter approved resume templates, layouts and formats. Choose your favorite and make it your own in minutes.

Free resume templates

Free resume builder

Improve your existing resume or start from scratch and create a standout, ATS-friendly resume. Add job-specific content, download and apply.

Free resume builder