Site Reliability Engineering (SRE) Manager

M&T Bank Corporation

Buffalo (NY)

On-site

USD 140,000 - 233,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

M&T Bank Corporation is seeking an experienced Site Reliability Engineering (SRE) Manager in Buffalo, NY to lead SRE and Production Support teams. The role focuses on reliability, observability, incident management, and automation across critical applications and cloud platforms.

You will drive the maturity of SLI/SLO governance, implement AI-enabled operational capabilities, and partner with engineering, security, and product stakeholders to ensure scalable, secure systems with optimal customer

Qualifications

  • 10+ years of technology experience with engineering, operations, or SRE.
  • 5+ years of leadership experience managing engineering or SRE teams.
  • Strong knowledge of SRE principles including SLOs, observability, automation, and incident management.
  • Experience with cloud platforms, APIs, and modern application architectures.

Responsibilities

  • Define and execute SRE strategies to improve reliability, availability, and performance.
  • Establish SLIs/SLOs and operational health metrics across services.
  • Lead major incident response and root cause analyses with corrective actions.
  • Drive automation and IaC to reduce toil and improve recovery times.
  • Promote AI-assisted operations and intelligent incident diagnostics.
  • Recruit, coach, and develop high-performing SRE teams and managers.

Skills

SRE Principles
Incident Management
Observability
Automation
Cloud Platforms
Leadership

Education

Bachelor's degree in Computer Science / Engineering

Tools

Dynatrace
Splunk
Datadog
Grafana
Azure Monitor
OpenTelemetry
PowerShell
Python
Bash

Job description

Overview

The Site Reliability Engineering (SRE) Manager leads teams responsible for the reliability, availability, performance, and operational excellence of critical business applications and platforms. This role combines engineering leadership with deep expertise in production operations, observability, automation, incident management, and cloud technologies. The SRE Manager partners with Engineering, Architecture, Infrastructure, Security, Product, and Business stakeholders to ensure systems are resilient, scalable, secure, and supportable. The role is accountable for driving operational excellence through automation, reliability engineering practices, and continuous improvement while developing high-performing SRE and Production Support teams.

Primary Responsibilities
Reliability & Operational Excellence
  • Define and execute SRE strategies that improve system reliability, availability, scalability, and performance.
  • Establish and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational health metrics.
  • Lead production readiness reviews, disaster recovery testing, resilience assessments, and operational risk mitigation activities.
  • Drive continuous improvement of application stability, service availability, and customer experience.
Incident & Problem Management
  • Lead major incident response and escalation management for critical production issues.
  • Oversee root cause analysis (RCA) processes and ensure corrective actions are implemented and tracked to completion.
  • Drive reduction of recurring incidents through engineering improvements, automation, and proactive monitoring.
  • Provide executive-level communication during significant incidents and service disruptions.
Observability & Automation
  • Establish monitoring, alerting, logging, tracing, and observability standards across supported platforms.
  • Lead implementation of dashboards and operational metrics that provide visibility into service health and customer impact.
  • Drive automation initiatives that reduce manual operational effort, improve recovery times, and increase engineering efficiency.
  • Promote Infrastructure as Code (IaC), CI/CD integration, automated remediation, and self-service operational capabilities.
Cloud & Platform Reliability
  • Partner with Engineering and Infrastructure teams to support cloud-native and hybrid application environments.
  • Ensure applications are designed and operated using resilient, scalable, and supportable architectures.
  • Support modernization initiatives involving Azure cloud services, containers, APIs, microservices, and platform engineering practices.
  • Evaluate vendor platforms and third-party services to ensure reliability and operational readiness.
AI & Modern Operations
  • Drive adoption of AI and Generative AI capabilities to improve incident response, troubleshooting, observability, and operational efficiency.
  • Identify opportunities for intelligent automation, anomaly detection, automated diagnostics, and AI-assisted knowledge management.
  • Promote responsible AI adoption aligned with enterprise security, governance, and risk standards.
People Leadership
  • Recruit, develop, coach, and retain high‑performing Site Reliability Engineers, Production Engineers, Automation Engineers, and Observability Engineers.
  • Establish career paths, skill development plans, and succession strategies.
  • Foster a culture of ownership, accountability, innovation, collaboration, and continuous learning.
  • Manage staffing, performance management, compensation recommendations, and organizational development activities.
Risk & Governance
  • Ensure adherence to enterprise risk, cybersecurity, regulatory, and operational control standards.
  • Identify and escalated operational risks impacting critical services or customer experiences.
  • Support audits, regulatory reviews, disaster recovery exercises, and operational governance programs.
Scope of Responsibilities

Leads teams responsible for: Site Reliability Engineering (SRE) Production Support Observability Engineering Incident Management Operational Automation Cloud Reliability Platform Operations

Responsible for reliability and operational health across multiple applications, platforms, cloud services, and vendor-supported solutions. Supervisory Responsibilities Typically manages 10-20 direct and indirect reports including SRE Engineers, Production Engineers, Technical Leads, and Engineering Managers.

Education & Experience Required

10+ years of technology experience with application support, infrastructure, cloud, software engineering, or reliability engineering responsibilities. 5+ years of leadership experience managing engineering, operations, or SRE teams. Experience managing production systems supporting critical business functions. Strong knowledge of Site Reliability Engineering principles, including SLOs, observability, automation, incident management, and operational excellence. Experience leading major incident response, root cause analysis, and service restoration efforts. Experience with cloud platforms, distributed systems, APIs, and modern application architectures. Strong communication, analytical, decision‑making, and stakeholder management skills.

Preferred Qualifications
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or related field.
  • Experience leading SRE or Production Engineering organizations.
  • Experience with Azure cloud technologies and cloud‑native architectures.
  • Experience with observability platforms such as Dynatrace, Splunk, Datadog, Grafana, Azure Monitor, or OpenTelemetry.
  • Experience with scripting and automation technologies including PowerShell, Python, Bash, and APIs.
  • Experience with CI/CD, Infrastructure as Code, DevOps, and Platform Engineering practices.
  • Experience implementing operational AI use cases including incident analysis, observability analytics, and automated diagnostics.
  • Financial services or other highly regulated industry experience preferred.
What Great Looks Like

A successful SRE Manager at M&T: Delivers highly available and resilient customer‑facing platforms. Uses automation to eliminate operational toil and improve efficiency. Reduces mean time to detect (MTTD) and mean time to restore (MTTR). Establishes strong observability and operational intelligence capabilities. Builds a culture of reliability, accountability, and continuous improvement. Successfully integrates AI‑assisted operations and automation into support workflows. develops high‑performing teams that balance reliability, speed, risk management, and customer experience.

Compensation

M&T Bank is committed to fair, competitive, and market‑informed pay for our employees. The pay range for this position is $139,700.00 - $232,900.00 Annual (USD). The successful candidate’s particular combination of knowledge, skills, and experience will inform their specific compensation.

Location

Buffalo, New York, United States of America

Company Culture

Great companies have an enduring sense of purpose. At M&T, our purpose is a simple one: make a difference in people’s lives and uplift the communities we serve. M&T Bank Corporation is a financial holding company headquartered in Buffalo, New York. M&T’s affiliates offer advice, guidance, expertise and solutions across the entire financial spectrum, combining M&T Bank’s traditional banking services with the wealth management and institutional capabilities offered by Wilmington Trust. M&T Bank has a network of over 1,000 branches and 2,200 ATMs that span 12 states from Maine to Virginia and Washington, D.C. For more than 165 years, M&T has strived to take an active role in our communities and build long‑lasting relationships with our customers. We are a bank for communities—combining the capabilities of a large bank with the care of a locally focused institution. As an employer of choice, we are proud to offer competitive benefits ranging from medical and retirement to forty hours of paid volunteer time, each year. Our core values – integrity, ownership, collaboration, curiosity, and candor – drive the work we do. We seek to further build upon our record of success by bringing in top talent and fresh skill sets while continuing to support the growth and development of all our team members. View M&T’s Human Capital Report to learn more.

M&T Bank is unwavering when it comes to providing equal employment opportunities to all employees and applicants without regard to race, color, national origin, religion, ethnicity, sex, gender identity, age, disability, citizenship, pregnancy, veteran status, military status, marital status, sexual orientation, genetic information or any other characteristic protected under applicable federal, state or local laws.

M&T Bank Corporation has policies and procedures in place to promote a drug free workplace.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Quality Engineering Manager
Quality Engineering Manager

M&T Bank Corporation • Buffalo (NY)

On-site
USD 116,000 - 194,000
Medical benefits
Retirement benefits
Volunteer time
Technical Engineer (Production Liability Engineer)
Technical Engineer (Production Liability Engineer)

M&T Bank Corporation • Buffalo (NY)

On-site
USD 97,000 - 162,000
Medical benefits
Retirement plan
Paid volunteer time
Engineering Manager (NOTA)
Engineering Manager (NOTA)

M&T Bank Corporation • Buffalo (NY)

On-site
USD 140,000 - 233,000
Business Systems Team Leader
Business Systems Team Leader

M&T Bank Corporation • Village of Williamsville (NY)

On-site
USD 103,000 - 172,000
Medical benefits
Retirement plan
Volunteer time off
Head of Forward Deployed Engineering (AI)
Head of Forward Deployed Engineering (AI)

M&T Bank Corporation • Buffalo (NY)

On-site
USD 140,000 - 233,000
Medical benefits
Retirement plans
Volunteer time off
Senior Software Engineer - Tech 360
Senior Software Engineer - Tech 360

M&T Bank Corporation • Buffalo (NY)

Hybrid
USD 97,000 - 162,000
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

M&T Bank • Buffalo (NY)

On-site
USD 140,000 - 233,000
Strategic Initiatives Lead
Strategic Initiatives Lead

M&T Bank Corporation • Buffalo (NY)

On-site
USD 90,000 - 149,000
Lead Software Engineer - Tech 360
Lead Software Engineer - Tech 360

M&T Bank Corporation • Buffalo (NY)

Hybrid
USD 116,000 - 194,000
CAM Manager IV
CAM Manager IV

M&T Bank Corporation • New York (NY), Northern (KY)

On-site
USD 108,000 - 179,000