Principal Site Reliability Engineer

M&T Bank

Buffalo (NY)

On-site

USD 140,000 - 233,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

M&T Bank is seeking a Senior Site Reliability Engineer to design, implement, and continuously improve scalable platform solutions. You will act as an SME in reliability engineering, driving standards, incident management, and automation across the Software Development Lifecycle.

Responsibilities include leading post-incident reviews, improving observability, and mentoring engineers. Strong collaboration with cross-functional teams and stakeholders is required.

Qualifications

  • Expert experience in system design, reliability engineering, and production operations.
  • Strong knowledge of SLOs/SLAs and incident management practices.
  • Experience with observability tooling and production readiness.
  • Cloud experience with AWS or Azure and CI/CD/DevOps practices.

Responsibilities

  • Define and drive service reliability standards across platforms.
  • Design highly available architectures aligned with scalability requirements.
  • Lead incident management, including detection, response, and recovery.
  • Drive root-cause analysis to prevent systemic issues.
  • Develop observability strategies: logging, monitoring, alerting, tracing.
  • Lead automation initiatives for self-healing systems and workflows.
  • Review roadmaps with reliability and performance considerations.
  • Partner with development, infrastructure, and architecture teams.
  • Serve as technical authority for performance and capacity planning.

Skills

SRE practices
Reliability engineering
Automation
Observability
Communication with stakeholders

Education

Associate’s degree
Bachelor’s degree

Tools

AWS
Azure
CI/CD tooling

Job description

Overview:

Responsible for designing, implementing, and continuously improving highly reliable, scalable, and resilient platform solutions across the enterprise. Operates as a subject matter expert (SME) in Site Reliability Engineering, driving reliability engineering practices, operational excellence, and automation across the Software Development Lifecycle. Leads complex initiatives, influences enterprise engineering standards, and partners with senior stakeholders to improve system stability, observability, and performance. Serves as a mentor and technical leader for less experienced engineers across Technology.

Primary Responsibilities:

  • Accountable for defining and driving service reliability standards, including SLOs, SLAs, and error budgets across platforms.
  • Design and implement highly available, fault-tolerant architectures aligned with enterprise scalability and resiliency requirements.
  • Lead incident management practices, including detection, response, escalation, and recovery processes.
  • Drive problem management and root cause analysis to prevent systemic issues.
  • Develop and promote observability strategies, including logging, monitoring, alerting, and tracing.
  • Lead automation initiatives for self-healing systems and operational workflows.
  • Contribute to and review technical roadmaps with reliability and performance considerations.
  • Partner with development, infrastructure, cybersecurity, and architecture teams.
  • Serve as a technical authority for performance, resilience, and capacity planning.
  • Drive production readiness practices including performance testing and failover capabilities.
  • Lead cross-team reliability improvement initiatives.
  • Participate in and lead post-incident reviews ensuring actionable outcomes.
  • Mentor engineers on reliability engineering and best practices.
  • Engage with stakeholders to identify risks and optimization opportunities.
  • Ensure adherence to risk and regulatory standards and elevate issues when needed.
  • Maintain internal control standards and compliance expectations.

Scope of Responsibilities:

Applies expert-level SRE practices across multiple platforms. Drives enterprise-wide reliability improvements and influences technical direction without direct authority.

Supervisory/Managerial Responsibilities:

No supervisory responsibilities.

Education and Experience Required:

Associate’s degree and a minimum of 9 years’ systems analysis and/ or application development work experience orBachelor’sdegree and a minimum of 7 years’ systems analysis and/ or application development work experience. In lieu of a degree, a combined minimum of 11 years’ education and/or relevant work experience, including a minimum of 7 years’ systems analysis and/ or application development work experience.

Expert experience in system design, reliability engineering, and production operations.

Advanced proficiency in at least one programming or scripting language.

Education and Experience Preferred:

Experience with observability and incident management tooling.

Experience with cloud platforms such as AWS or Azure.

Strong understanding of CI/CD, DevOps, and SDLC practices.

Experience defining and implementing SLO/SLI frameworks.

Experience in regulated environments such as financial services.

Strong communication and stakeholder management skills.

M&T Bank is committed to fair, competitive, and market-informed pay for our employees. The pay range for this position is $139,700.00 - $232,900.00 Annual (USD). The successful candidate’s particular combination of knowledge, skills, and experience will inform their specific compensation.

Location

Buffalo, New York, United States of America

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

M&T Bank • Buffalo (NY)

On-site
USD 140,000 - 233,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hobbsnews • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Discretionary incentive plan
Senior Site Reliability Engineer
Senior Site Reliability Engineer

National Black MBA Association • Jersey City (NJ)

On-site
USD 153,000 - 192,000
Benefits eligible
Annual discretionary plan
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Luxoft • United States

On-site
USD 140,000 - 180,000
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

M&T Bank Corporation • Buffalo (NY)

On-site
USD 140,000 - 233,000
Medical benefits
Retirement plan
Volunteer time off
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Jobgether • United States

Hybrid
USD 120,000 - 160,000
Competitive compensation package
Flexible work arrangements
Professional development opportunities
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Koitecc Solutions • Plano (TX), Northern (KY)

Hybrid
USD 153,000 - 192,000
Discretionary incentive
Benefits package
Senior Site Reliability Engineer - Banking & Finance
Senior Site Reliability Engineer - Banking & Finance

Hamilton Barnes Associates Limited • New York (NY)

Hybrid
USD 360,000 - 440,000
Strong compensation and bonus potential
Collaborative engineering culture
Work on mission-critical systems
Lead Site Reliability Engineer
Lead Site Reliability Engineer

CardWorks Servicing LLC • Pittsburgh

On-site
USD 146,000 - 163,000
Competitive base pay
Medical, Dental and Vision coverage
401(k) Plan with Company Match
+1
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Weekday (YC W21) • New York (NY)

On-site
USD 150,000 - 250,000
Health, dental, vision insurance
Generous PTO
Learning & development
+2