Senior Site Reliability Engineer

ISO New England

Manchester (NH)

Hybrid

USD 134,000 - 170,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Performance bonus
401(k) with employer contributions
Tuition reimbursement & professional发展
Wellness programs & onsite gym
Hybrid work environment (3 days onsite
Relocation assistance
Free coffee at onsite cafe

Job summary

ISO New England is seeking a Senior Site Reliability Engineer to improve reliability, observability, and automation across IT services. You will implement scalable IaC, manage Splunk platforms, and reduce toil with automation.

The role emphasizes observability, Terraform/IaC, Python/PowerShell, and AI-assisted development tools to accelerate solutions while ensuring security and reliability. Hybrid work is offered in a mission-focused environment.

Qualifications

  • 5 years of experience in SRE, DevOps, systems engineering, platform engineering, or IT operations
  • Experience with enterprise monitoring and observability platforms. Hands-on experience with Splunk is strongly preferred. Candidates without direct Splunk experience must demonstrate a strong willingness to develop expertise in Splunk administration, engineering, and automation
  • Experience designing, deploying, or managing infrastructure using Terraform and Infrastructure as Code (IaC) practices
  • Strong scripting and automation experience using Python, PowerShell, Bash, or similar technologies, including development of operational tooling and workflow automation in production environments
  • Demonstrated experience designing, developing, and supporting automation solutions that reduced manual operational effort in an enterprise environment
  • Ability to read, understand, review, troubleshoot, and refine code produced by engineering teams or AI-assisted development platforms
  • Knowledge of distributed systems, networking, enterprise infrastructure, and cloud platforms
  • Familiarity with SRE principles including SLOs, error budgets, observability, and toil reduction
  • Ability to analyze and troubleshoot complex technical systems

Responsibilities

  • Build and maintain observability, monitoring, logging, alerting, and telemetry platforms (e.g., Splunk, Dynatrace, PRTG, OpsGenie, StatusPage)
  • Administer, maintain, automate, and improve the Splunk platform—data onboarding, indexing, search performance, dashboards, access controls, health monitoring, scalability
  • Develop and automate Splunk onboarding, configuration, monitoring, and operational workflows
  • Develop meaningful KPIs and dashboards for IT service health
  • Engineer resilience patterns including HA, DR, and automated failover
  • Plan and execute resilience testing and failover exercises with infrastructure and application teams
  • Perform performance testing, capacity modeling, forecasting, and right-sizing
  • Participate in major incident response to accelerate service restoration
  • Identify and eliminate manual operational toil through automation
  • Design, deploy, and manage infrastructure using Terraform and IaC practices
  • Identify gaps in observability coverage and drive solutions
  • Collaborate with architecture and application teams for production readiness
  • Leverage AI-assisted development tools to accelerate automation while ensuring reliability
  • Reduce repeat incidents by engineering permanent fixes and continuous improvement

Skills

Splunk
Terraform
Python
PowerShell
Bash
Automation
Observability
SRE Principles
CI/CD

Tools

Dynatrace
PRTG
OpsGenie
StatusPage

Job description

ISO New England is the independent system operator responsible for ensuring the safe and reliable flow of electricity in our region and planning for the future of the electric grid. We are at the forefront of New England's ongoing transition to clean energy.

The Senior Site Reliability Engineer (SRE) is a hands-on engineering role responsible for improving the reliability, observability, performance, and operational efficiency of ISO New England's IT services. The SRE works across infrastructure, platform, cyber security, and application teams to reduce operational toil, improve service resilience, and implement scalable automation solutions.

This role has a strong emphasis on observability engineering, automation, Splunk administration, and Infrastructure as Code (IaC). The ideal candidate will possess hands-on experience with Splunk or demonstrate a strong willingness to develop expertise in the platform. Experience with Terraform, automation technologies such as Python and PowerShell, and the ability to leverage AI-assisted development tools to accelerate engineering solutions are key components of the role.

What we offer you:
  • A stable, mission-driven workplace where your impact truly matters

  • A highly engaged work environment that values inclusion, collaboration, and employee safety and wellbeing

  • Competitive compensation with a base salary performance bonus

  • Robust benefits package, including:

  • Enhanced 401(k) and financial planning support

  • Tuition reimbursement and professional development

  • Wellness programs, including an onsite gym

  • Flexible work hours

  • Employee Business Networks

  • Free coffee at our onsite caf

  • Hybrid work environment (3 days/week onsite)

  • Distance-based relocation assistance available

How you will make an Impact
  • Build and maintain observability, monitoring, logging, alerting, and telemetry platforms (e.g., Splunk, Dynatrace, PRTG, OpsGenie, StatusPage)
  • Administer, maintain, automate, and continuously improve the Splunk platform, including data onboarding, indexing, search performance, dashboards, access controls, health monitoring, platform scalability, and operational workflows
  • Develop and automate Splunk onboarding, configuration, monitoring, and operational workflows to improve platform reliability and reduce administrative overhead
  • Develop meaningful KPIs and dashboards for business and IT service health
  • Engineer and implement resilience patterns including HA, DR, and automated failover
  • Partner with infrastructure and application teams to plan and execute resilience testing and failover exercises to validate recovery capabilities and observability coverage
  • Conduct performance testing, capacity modeling, forecasting, and right-sizing
  • Participate in major incident response activities, providing technical expertise to accelerate service restoration and identify reliability improvements
  • Identify, prioritize, and eliminate manual operational toil through automation, targeting workflows, runbooks, alerting, platform administration, service management processes, and KPI collection, with a bias toward scalable and repeatable engineering solutions
  • Design, develop, maintain, and support automation solutions, integrations, and operational tooling using Python, PowerShell, Bash, or similar technologies to improve reliability, reduce manual effort, and enhance operational efficiency
  • Design, deploy, and manage infrastructure using Terraform and Infrastructure as Code (IaC) practices, including observability platforms, infrastructure services, and supporting technology stacks, with a focus on consistency, repeatability, and operational sustainability
  • Identify gaps in observability coverage and drive engineering solutions to close them
  • Collaborate with architecture and application teams to ensure production readiness
  • Leverage AI-assisted development tools to accelerate automation initiatives while reviewing, validating, troubleshooting, and refining generated code to ensure reliability, security, maintainability, and operational effectiveness
  • Reduce repeat incidents by engineering permanent fixes and driving continuous improvement
What we are looking for
  • 5 years of experience in SRE, DevOps, systems engineering, platform engineering, or IT operations
  • Experience with enterprise monitoring and observability platforms. Hands-on experience with Splunk is strongly preferred. Candidates without direct Splunk experience must demonstrate a strong willingness and aptitude to develop expertise in Splunk administration, engineering, and automation.
  • Experience designing, deploying, or managing infrastructure using Terraform and Infrastructure as Code (IaC) practices
  • Strong scripting and automation experience using Python, PowerShell, Bash, or similar technologies, including the development of operational tooling, integrations, and workflow automation in production environments
  • Demonstrated experience designing, developing, and supporting automation solutions that measurably reduced manual operational effort in an enterprise environment
  • Ability to read, understand, review, troubleshoot, and refine code produced by engineering teams or AI-assisted development platforms
  • Knowledge of distributed systems, networking, enterprise infrastructure, and cloud platforms
  • Familiarity with SRE principles including SLOs, error budgets, observability, and toil reduction
  • Ability to analyze and troubleshoot complex technical systems
Preferred Qualifications
  • Experience in mission-critical, highly available, or regulated environments
  • Experience utilizing AI-assisted development tools to accelerate automation, operational engineering, or platform management activities
  • Knowledge of ITIL processes and/or SRE best practices
  • Experience with performance testing, capacity planning, resilience testing, or disaster recovery validation

This employer will not sponsor applicants for work visas for this position (ex: H-1B, F-1/CPT/OPT, O-1, E-3, TN, J, etc.).

The expected salary range for this position is $134,000 - $170,000 per year, for a Senior to Lead level candidate. This role is also eligible for an annual performance bonus, comprehensive health insurance (medical, dental and vision), flexible spending and health savings accounts, a 401(k) plan with generous employer contributions and a student debt benefit, life and AD&D insurance, disability insurance, critical illness and hospital indemnity benefits, paid time off, paid leave, a wellness program, an employee assistance program and other great company perks.

#LI-HYBRID

This is a U.S. based role. If the successful candidate resides outside of the U.S., relocation will be required.

Equal Opportunity

We are proud to be an EEO employer. Applicants for employment are considered without regard to race, color, religion, creed, sex (including pregnancy, childbirth, and related medical conditions), gender identity or expression, sexual orientation, citizenship, national origin, age, ancestry, marital status, disability (including learning, mental, intellectual, and physical), service in the uniformed services, genetic information, or any other status protected by applicable law.

Drug Free Environment

We maintain a drug-free workplace and perform pre-employment substance abuse testing.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

ISO New England Inc. • Holyoke (MA)

Hybrid
USD 134,000 - 170,000
Hybrid work environment (3 days/week)
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Chicago (IL)

Hybrid
USD 140,000 - 170,000
Comprehensive healthcare benefits
401(k) match
Paid Time Off (PTO)
+2
R&D Analyst/Engineer
R&D Analyst/Engineer

ISO New England Inc. • Windsor (CT)

Hybrid
USD 100,000 - 135,000
Performance bonus
Wellness programs
Tuition reimbursement
+2
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

SEI • Oaks (PA)

Hybrid
USD 140,000 - 170,000
Comprehensive healthcare coverage
401(k) matching
Tuition reimbursement
+1
Senior DevOps/SRE Engineer
Senior DevOps/SRE Engineer

VITG • Ellicott City (MD)

Hybrid
USD 90,000 - 120,000
401(k) with employer contribution
Medical/Dental/Vision insurance
Paid vacation (PTO)
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Bellevue (CA)

On-site
USD 180,000 - 230,000
Amazing Benefits
Making Social Impact
Fostering Diversity, Equity, Inclusion
Senior SRE: Observability & Automation Lead
Senior SRE: Observability & Automation Lead

ISO New England Inc. • Holyoke (MA)

Hybrid
USD 134,000 - 170,000
Enhanced 401(k) and financial planning
Tuition reimbursement and professional
Onsite gym and wellness programs
+4
Staff TDI Site Reliability Engineer, Okta Federal
Staff TDI Site Reliability Engineer, Okta Federal

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 174,000 - 239,000
Staff Site Reliability Engineer - Splunk
Staff Site Reliability Engineer - Splunk

United States Digital Space LLC • Washington

On-site
USD 194,000 - 267,000
Senior Manager, Site Reliability Engineering - Infrastructure Platform
Senior Manager, Site Reliability Engineering - Infrastructure Platform

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 232,000 - 319,000
Equity
Bonus
Health insurance
+2