Technical Engineer (Production Liability Engineer)

M&T Bank Corporation

Buffalo (NY)

On-site

USD 97,000 - 162,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical benefits
Retirement plan
Paid volunteer time

Job summary

M&T Bank Corporation in Buffalo, NY is seeking a Technical Engineer to lead production support, incident management, and SRE practices for critical applications and platforms.

You will champion observability, automation, IaC, and AI-assisted operations, partnering across engineering, security, product, and vendor teams to improve reliability and resilience.

Qualifications

  • Experience with Site Reliability Engineering (SRE) practices.
  • Experience defining and managing SLIs, SLOs, and operational metrics.
  • Experience supporting distributed systems, APIs, microservices, and cloud-native applications.

Responsibilities

  • Serve as a technical escalation point for critical production incidents and outages.
  • Lead troubleshooting and RCA activities across applications, infrastructure, cloud, and integrations.
  • Coordinate incident response with engineering, infrastructure, vendors, and business stakeholders.

Skills

SRE practices
Incident management
Observability
Automation
IaC
Troubleshooting
Cloud operations
AI-assisted operations

Tools

Dynatrace
Splunk
Datadog
Grafana
Prometheus
Azure Monitor

Job description

Overview

The Technical Engineer serves as a senior production support and reliability engineering professional responsible for ensuring the availability, stability, performance, and operational excellence of critical business applications and platforms. This role combines strong troubleshooting expertise with modern Site Reliability Engineering (SRE), observability, automation, cloud operations, and incident management practices. The Technical Engineer partners with Engineering, Architecture, Infrastructure, Security, Product, and Vendor teams to proactively identify operational risks, improve system resilience, accelerate incident resolution, and continuously enhance customer and employee experiences. The ideal candidate possesses deep technical knowledge of application support, distributed systems, cloud technologies, monitoring platforms, automation tools, and modern operational practices. They are passionate about eliminating repetitive work through automation and leveraging AI-powered tools to improve operational efficiency and support outcomes.

Primary Responsibilities
Production Support & Incident Management

Serve as a technical escalation point for critical production incidents, outages, and service degradation events. Lead troubleshooting and root cause analysis efforts across applications, integrations, infrastructure, cloud services, APIs, and supporting technologies. Coordinate incident response activities involving application teams, infrastructure teams, vendors, and business stakeholders. Restore service quickly while ensuring long-term corrective actions are identified and implemented. Participate in major incident management processes and post-incident reviews.

Site Reliability Engineering (SRE)

Apply Site Reliability Engineering principles to improve platform reliability, scalability, resilience, and operational efficiency. Define and support Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational performance metrics. Drive reduction of operational toil through automation and process improvement. Support production readiness reviews and operational acceptance processes. Participate in disaster recovery, resiliency, failover, and business continuity testing.

Troubleshooting & Problem Management

Analyze complex system behavior using logs, metrics, traces, performance data, and monitoring tools. Perform deep technical investigations across application, infrastructure, data, network, and cloud environments. Identify recurring issues, trends, and systemic problems to reduce future incidents. Lead root cause analysis (RCA) activities and implement preventive solutions. Develop technical recommendations that improve system stability, performance, and reliability.

Observability & Monitoring

Design, implement, and optimize monitoring, alerting, logging, and observability solutions. Develop dashboards and health indicators providing visibility into application and platform performance. Partner with engineering teams to improve observability through instrumentation, distributed tracing, synthetic monitoring, and telemetry collection. Continuously refine alerting strategies to reduce false positives and alert fatigue. Establish operational health metrics and reliability reporting.

Automation & Scripting

Develop and maintain automation solutions that improve operational efficiency and service reliability. Create scripts, tools, and workflows to automate diagnostics, health checks, remediation activities, and routine support tasks. Leverage Infrastructure as Code (IaC) and automation frameworks where appropriate. Drive continuous improvement through operational automation and self-healing capabilities. Partner with engineering teams to integrate automation into deployment and operational workflows.

Desired Scripting Technologies
  • PowerShell
  • Python
  • Bash/Shell
  • SQL
  • REST APIs
  • Workflow automation platforms
  • AI-Assisted Operations & Innovation
AI-Assisted Operations & Innovation

Leverage AI and Generative AI tools to improve incident analysis, troubleshooting, knowledge management, and operational efficiency. Utilize AI-powered operational insights to identify patterns, anomalies, and emerging risks. Contribute to development of intelligent support capabilities including chatbots, operational copilots, automated RCA generation, and knowledge recommendations. Evaluate opportunities to improve production support through AI-enabled automation and predictive analytics. Promote responsible AI practices aligned with enterprise governance and security requirements.

Support Playbooks & Knowledge Management

Develop, maintain, and continuously improve support runbooks, operational procedures, troubleshooting guides, and recovery playbooks. Ensure support documentation remains accurate, actionable, and aligned with production environments. Establish standardized operational processes supporting incident response and service recovery. Capture lessons learned from incidents and incorporate improvements into support practices. Build and maintain operational knowledge repositories to improve support consistency and reduce resolution times.

Cloud Operations

Support cloud-hosted and hybrid application environments, including Azure-based platforms and services. Assist engineering teams in implementing resilient and observable cloud architectures. Monitor cloud resource health, performance, utilization, and operational readiness. Support cloud deployments, platform upgrades, and operational change activities. Partner with cloud engineering teams on modernization and reliability initiatives.

Collaboration & Technical Leadership

Partner closely with Engineering, Product, Architecture, Infrastructure, Security, QA, and Vendor teams. Review operational readiness of new systems and platform enhancements. Mentor junior support engineers and provide technical guidance during incident response activities. Promote operational excellence, reliability engineering principles, and continuous improvement practices. Serve as a subject matter expert within assigned technology domains.

Governance, Risk & Compliance

Ensure production support activities comply with enterprise risk, security, regulatory, and audit requirements. Identify operational risks and escape issues appropriately. Support implementation of internal controls and operational governance standards. Participate in audit, compliance, and regulatory review activities as required.

Preferred Qualifications
Reliability Engineering & Operations
  • Experience with Site Reliability Engineering (SRE) practices.
  • Experience defining and managing SLIs, SLOs, and operational metrics.
  • Experience supporting distributed systems, APIs, microservices, and cloud-native applications.
  • Experience performing production readiness reviews and operational assessments.
Observability & Monitoring
  • Experience with platforms such as: Dynatrace Splunk Datadog Azure Monitor Grafana Prometheus OpenTelemetry AppDynamics Cloud Technologies.
  • Experience supporting Azure cloud environments.
  • Familiarity with Azure App Services, AKS, Functions, Storage, Event Hub, Service Bus, and monitoring services.
  • Understanding of cloud security and operational best practices.
Automation & Scripting
  • Advanced PowerShell scripting.
  • Python development and automation.
  • REST API integration.
  • Infrastructure as Code concepts.
  • CI/CD tools and deployment pipelines.
AI & Modern Operations
  • Experience using Microsoft Copilot, Azure AI services, Microsoft Foundry, or similar AI-enabled platforms.
  • Familiarity with AI-assisted troubleshooting and operational analytics.
  • Experience implementing AI-enabled support workflows or knowledge management solutions.
  • Understanding of intelligent automation and operational copilots.
What Great Looks Like
  • Resolves complex incidents quickly and effectively.
  • Proactively identifies and eliminates recurring production issues.
  • Uses automation and scripting to eliminate manual support work.
  • Builds comprehensive support playbooks and operational runbooks.
  • Leverages AI to accelerate troubleshooting and operational insights.
  • Establishes strong observability and monitoring capabilities.
  • Partners effectively with engineering teams to improve reliability.
  • Drives a culture of operational excellence through SRE principles.
  • Continuously improves customer experience through system stability, resilience, and performance.
M&T Bank is committed to fair, competitive, and market-informed pay for our employees.

The pay range for this position is $97,100.00 - $161,800.00 Annual (USD).

Location

Buffalo, New York, United States of America

Benefits
  • Competitive benefits ranging from medical and retirement to forty hours of paid volunteer time, each year.
Company Values

Our core values – integrity, ownership, collaboration, curiosity, and candor – drive the work we do.

We seek to further build upon our record of success by bringing in top talent and fresh skill sets while continuing to support the growth and development of all our team members.

View M&T’s Human Capital Report to learn more.

M&T Bank Corporation has policies and procedures in place to promote a drug free workplace.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Quality Engineering Manager
Quality Engineering Manager

M&T Bank Corporation • Buffalo (NY)

On-site
USD 116,000 - 194,000
Medical benefits
Retirement benefits
Volunteer time
Site Reliability Engineering (SRE) Manager
Site Reliability Engineering (SRE) Manager

M&T Bank Corporation • Buffalo (NY)

On-site
USD 140,000 - 233,000
Quality Engineering Manager
Quality Engineering Manager

M&T Bank • Buffalo (NY)

On-site
USD 116,000 - 194,000
Technical Engineer (Production Liability Engineer)
Technical Engineer (Production Liability Engineer)

M&T Bank • Buffalo (NY)

On-site
USD 97,000 - 162,000
Engineering Manager (NOTA)
Engineering Manager (NOTA)

M&T Bank Corporation • Buffalo (NY)

On-site
USD 140,000 - 233,000
Lead Software Engineer - Tech 360
Lead Software Engineer - Tech 360

M&T Bank Corporation • Buffalo (NY)

Hybrid
USD 116,000 - 194,000
Technology Senior Manager - Service Reliability & Operations
Technology Senior Manager - Service Reliability & Operations

M&T Bank • Buffalo (NY)

On-site
USD 148,000 - 247,000
Competitive compensation
Health, welfare, and retirement
25 days PTO + 12 holidays
+2
Business Systems Team Leader
Business Systems Team Leader

M&T Bank Corporation • Village of Williamsville (NY)

On-site
USD 103,000 - 172,000
Medical benefits
Retirement plan
Volunteer time off
Senior Software Engineer - Tech 360
Senior Software Engineer - Tech 360

M&T Bank Corporation • Buffalo (NY)

Hybrid
USD 97,000 - 162,000
Manager, Site Reliability Engineering
Manager, Site Reliability Engineering

M&T Bank • Buffalo (NY)

On-site
USD 140,000 - 233,000