Site Reliability Engineering (SRE) Manager

mtb

Buffalo (NY)

Hybrid

USD 150,000 - 230,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

mtb seeks an experienced Site Reliability Engineering (SRE) Manager to lead multiple teams responsible for reliability, performance, and operational excellence across critical applications. You will drive incident response, RCA processes, and continuous improvement of monitoring, automation, and platform health.

The role requires 10+ years in technology with leadership experience, strong SRE fundamentals, and collaboration with cloud, security, and product stakeholders.

Qualifications

  • 10+ years in technology roles with reliability responsibilities.
  • Proven leadership of SRE or production teams.
  • Strong knowledge of SRE principles, SLIs/SLOs, incident response.

Responsibilities

  • Define and execute SRE strategies to improve reliability and performance.
  • Lead incident response, RCA, and post-mortem actions.
  • Drive observability standards, dashboards, and metrics.
  • Promote IaC, CI/CD, automation, and self-service platforms.
  • Collaborate with cloud and platform teams on reliability initiatives.

Skills

SRE leadership
Observability
Automation
Incident management
Cloud platforms

Tools

IaC
CI/CD
Azure
Containers

Job description

Overview

The Site Reliability Engineering (SRE) Manager leads teams responsible for the reliability, availability, performance, and operational excellence of critical business applications and platforms. This role combines engineering leadership with deep expertise in production operations, observability, automation, incident management, and cloud technologies.

The SRE Manager partners with Engineering, Architecture, Infrastructure, Security, Product, and Business stakeholders to ensure systems are resilient, scalable, secure, and supportable. The role is accountable for driving operational excellence through automation, reliability engineering practices, and continuous improvement while developing high-performing SRE and Production Support teams.

Primary Responsibilities
Reliability & Operational Excellence

Define and execute SRE strategies that improve system reliability, availability, scalability, and performance.

Establish and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational health metrics.

Lead production readiness reviews, disaster recovery testing, resilience assessments, and operational risk mitigation activities.

Drive continuous improvement of application stability, service availability, and customer experience.

Incident & Problem Management

Lead major incident response and escalation management for critical production issues.

Oversee root cause analysis (RCA) processes and ensure corrective actions are implemented and tracked to completion.

Drive reduction of recurring incidents through engineering improvements, automation, and proactive monitoring.

Provide executive-level communication during significant incidents and service disruptions.

Observability & Automation

Establish monitoring, alerting, logging, tracing, and observability standards across supported platforms.

Lead implementation of dashboards and operational metrics that provide visibility into service health and customer impact.

Drive automation initiatives that reduce manual operational effort, improve recovery times, and increase engineering efficiency.

Promote Infrastructure as Code (IaC), CI/CD integration, automated remediation, and self-service operational capabilities.

Cloud & Platform Reliability

Partner with Engineering and Infrastructure teams to support cloud-native and hybrid application environments.

Ensure applications are designed and operated using resilient, scalable, and supportable architectures.

Support modernization initiatives involving Azure cloud services, containers, APIs, microservices, and platform engineering practices.

Evaluate vendor platforms and third-party services to ensure reliability and operational readiness.

AI & Modern Operations

Drive adoption of AI and Generative AI capabilities to improve incident response, troubleshooting, observability, and operational efficiency.

Identify opportunities for intelligent automation, anomaly detection, automated diagnostics, and AI-assisted knowledge management.

Promote responsible AI adoption aligned with enterprise security, governance, and risk standards.

People Leadership

Recruit, develop, coach, and retain high-performing Site Reliability Engineers, Production Engineers, Automation Engineers, and Observability Engineers.

Establish career paths, skill development plans, and succession strategies.

Foster a culture of ownership, accountability, innovation, collaboration, and continuous learning.

Manage staffing, performance management, compensation recommendations, and organizational development activities.

Risk & Governance

Ensure adherence to enterprise risk, cybersecurity, regulatory, and operational control standards.

Identify and elevate operational risks impacting critical services or customer experiences.

Support audits, regulatory reviews, disaster recovery exercises, and operational governance programs.

Scope of Responsibilities

Leads teams responsible for:

  • Site Reliability Engineering (SRE)
  • Production Support
  • Observability Engineering
  • Incident Management
  • Operational Automation
  • Cloud Reliability
  • Platform Operations

Responsible for reliability and operational health across multiple applications, platforms, cloud services, and vendor-supported solutions.

Supervisory Responsibilities

Typically manages 10-20 direct and indirect reports including SRE Engineers, Production Engineers, Technical Leads, and Engineering Managers.

Education & Experience Required

10+ years of technology experience with application support, infrastructure, cloud, software engineering, or reliability engineering responsibilities.

5+ years of leadership experience managing engineering, operations, or SRE teams.

Experience managing production systems supporting critical business functions.

Strong knowledge of Site Reliability Engineering principles, including SLOs, observability, automation, incident management, and operational excellence.

Experience leading major incident response, root cause analysis, and service restoration efforts.

E

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Manager SRE
Senior Manager SRE

Expedite Talent Solutions • United States

Hybrid
USD 130,000 - 160,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

On-site
USD 130,000 - 180,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Harvey Nash • Charlotte (NC)

On-site
USD 100,000 - 130,000
Full-Time Lead Site Reliability Engineer
Full-Time Lead Site Reliability Engineer

TSP talent • O’Fallon (MO)

On-site
USD 120,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Charlotte (TX)

On-site
USD 153,000 - 192,000
Discretionary incentive eligible
Benefits package
Quality Engineering Manager
Quality Engineering Manager

mtb • Buffalo (NY)

On-site
USD 90,000 - 130,000
SRE Manager: Reliability, Automation & Platform Ops
SRE Manager: Reliability, Automation & Platform Ops

mtb • Buffalo (NY)

Hybrid
USD 150,000 - 230,000
Site Reliability Engineer
Site Reliability Engineer

IntraEdge • Austin (TX)

On-site
USD 120,000 - 180,000