Technical Engineer (Production Liability Engineer)

mtb

Buffalo (NY)

On-site

USD 110,000 - 150,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

mtb is seeking a Technical Engineer to lead production support and reliability engineering for critical business applications. You will apply SRE practices, observability, automation, and cloud operations to improve resilience, reduce toil, and accelerate incident resolution.

You will partner with engineering, security, product, and vendor teams to design proactive reliability strategies, implement automated diagnostics, and evolve runbooks and knowledge resources for rapid recovery and stable

Qualifications

  • Senior production support and reliability engineering experience.
  • Experience with observability, automation, cloud operations, and incident management.
  • Familiarity with AI-assisted operations.
  • Ability to reduce toil and improve platform reliability.

Responsibilities

  • Serve as technical escalation point for critical incidents.
  • Lead troubleshooting and root cause analysis across applications and infrastructure.
  • Coordinate incident response with multiple teams.
  • Restore service quickly and implement long-term fixes.
  • Participate in disaster recovery and business continuity testing.
  • Design and optimize monitoring, alerting, and observability.
  • Develop automation scripts and IaC to improve reliability.
  • Create runbooks and knowledge management content.

Skills

SRE
Observability
Automation
Incident Mgmt
Troubleshooting
Cloud ops
AI Ops
Scripting
Monitoring

Job description

Overview

The Technical Engineer serves as a senior production support and reliability engineering professional responsible for ensuring the availability, stability, performance, and operational excellence of critical business applications and platforms.

This role combines strong troubleshooting expertise with modern Site Reliability Engineering (SRE), observability, automation, cloud operations, and incident management practices. The Technical Engineer partners with Engineering, Architecture, Infrastructure, Security, Product, and Vendor teams to proactively identify operational risks, improve system resilience, accelerate incident resolution, and continuously enhance customer and employee experiences.

The ideal candidate possesses deep technical knowledge of application support, distributed systems, cloud technologies, monitoring platforms, automation tools, and modern operational practices. They are passionate about eliminating repetitive work through automation and leveraging AI-powered tools to improve operational efficiency and support outcomes.

Primary Responsibilities
Production Support & Incident Management

Serve as a technical escalation point for critical production incidents, outages, and service degradation events.

Lead troubleshooting and root cause analysis efforts across applications, integrations, infrastructure, cloud services, APIs, and supporting technologies.

Coordinate incident response activities involving application teams, infrastructure teams, vendors, and business stakeholders.

Restore service quickly while ensuring long-term corrective actions are identified and implemented.

Participate in major incident management processes and post-incident reviews.

Site Reliability Engineering (SRE)

Apply Site Reliability Engineering principles to improve platform reliability, scalability, resilience, and operational efficiency.

Define and support Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational performance metrics.

Drive reduction of operational toil through automation and process improvement.

Support production readiness reviews and operational acceptance processes.

Participate in disaster recovery, resiliency, failover, and business continuity testing.

Troubleshooting & Problem Management

Analyze complex system behavior using logs, metrics, traces, performance data, and monitoring tools.

Perform deep technical investigations across application, infrastructure, data, network, and cloud environments.

Identify recurring issues, trends, and systemic problems to reduce future incidents.

Lead root cause analysis (RCA) activities and implement preventive solutions.

Develop technical recommendations that improve system stability, performance, and reliability.

Observability & Monitoring

Design, implement, and optimize monitoring, alerting, logging, and observability solutions.

Develop dashboards and health indicators providing visibility into application and platform performance.

Partner with engineering teams to improve observability through instrumentation, distributed tracing, synthetic monitoring, and telemetry collection.

Continuously refine alerting strategies to reduce false positives and alert fatigue.

Establish operational health metrics and reliability reporting.

Automation & Scripting

Develop and maintain automation solutions that improve operational efficiency and service reliability.

Create scripts, tools, and workflows to automate diagnostics, health checks, remediation activities, and routine support tasks.

Leverage Infrastructure as Code (IaC) and automation frameworks where appropriate.

Drive continuous improvement through operational automation and self-healing capabilities.

Partner with engineering teams to integrate automation into deployment and operational workflows.

Desired Scripting Technologies:
  • PowerShell
  • Python
  • Bash/Shell
  • SQL
  • REST APIs
  • Workflow automation platforms
AI-Assisted Operations & Innovation

Leverage AI and Generative AI tools to improve incident analysis, troubleshooting, knowledge management, and operational efficiency.

Utilize AI-powered operational insights to identify patterns, anomalies, and emerging risks.

Contribute to development of intelligent support capabilities including chatbots, operational copilots, automated RCA generation, and knowledge recommendations.

Evaluate opportunities to improve production support through AI-enabled automation and predictive analytics.

Promote responsible AI practices aligned with enterprise governance and security requirements.

Support Playbooks & Knowledge Management

Develop, maintain, and continuously improve support runbooks, operational procedures, troubleshooting guides, and recovery playbooks.

Ensure support documentation remains accurate, actionable, and aligned with production environments.

Establish standardized operational processes supporting incident response and service recovery.

Capture lessons learned from incidents and incorporate improvements into support practice

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Quality Engineering Manager
Quality Engineering Manager

mtb • Buffalo (NY)

On-site
USD 90,000 - 130,000
Technical Engineer (Production Liability Engineer)
Technical Engineer (Production Liability Engineer)

M&T Bank • Buffalo (NY)

On-site
USD 97,000 - 162,000
Quality Engineering Manager
Quality Engineering Manager

M&T Bank • Buffalo (NY)

On-site
USD 116,000 - 194,000
Application Engineer View role →
Application Engineer View role →

NRnP Technology • Northern (KY)

On-site
USD 90,000 - 140,000
Technical Engineer (Production Liability Engineer)
Technical Engineer (Production Liability Engineer)

M&T Bank • South Dakota

On-site
USD 97,000 - 162,000
Site Reliability Engineering (SRE) Manager
Site Reliability Engineering (SRE) Manager

mtb • Buffalo (NY)

Hybrid
USD 150,000 - 230,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Technology Support Lead - Problem Management & Governance Lead
Technology Support Lead - Problem Management & Governance Lead

JPMorgan Chase & Co. • Columbus (OH)

On-site
USD 120,000 - 180,000
Technical Lead-Cloud & Infra Engg
Technical Lead-Cloud & Infra Engg

Birlasoft • New Jersey

On-site
USD 140,000 - 190,000
Site Reliability Engineer (SRE) – Production Services
Site Reliability Engineer (SRE) – Production Services

TechDigital Group • Pittsburgh

On-site
USD 120,000 - 160,000