SRE Lead: Incident Commander & Reliability Champion

Us Bank

Irving (TX)

On-site

USD 112,000 - 131,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
401(k) plan
Paid vacation
12 holidays
Parental leave
Disability insurance

Job summary

U.S. Bank is seeking an experienced Site Reliability Engineer to lead complex incident resolution, drive reliability initiatives, and collaborate with software, infra, and product teams.

The role emphasizes RCA, monitoring, IaC, CI/CD, and proactive resilience across AWS/Azure/Kubernetes environments. Eligible candidates should have strong leadership and communication skills for cross-functional coordination.

Qualifications

  • Bachelor's degree, or equivalent work experience.
  • 6–8 years of relevant work experience in IT service management, production support, or related areas.

Responsibilities

  • Lead troubleshooting and resolution of complex production incidents.
  • Conduct root cause analysis (RCA) and implement permanent corrective actions.
  • Design and enhance monitoring, observability, and runbooks for reliability.
  • Drive automation via scripting, IaC, CI/CD, and self-healing capabilities.
  • Coordinate cross-functional incident response as Incident Commander.
  • Mentor and manage SRE, DevOps, and production support engineers.
  • Track MTTR/MTTD, SLA compliance, and backlog health to drive improvements.

Skills

SRE
DevOps
Production Support
Platform Engineering
Distributed Systems
Incident Management
Problem Management
Change Management
RCA
AWS
Azure
Kubernetes
Docker
CI/CD
GitHub Actions
Azure DevOps
Jenkins
GitLab
Datadog
Splunk
Dynatrace
Grafana
Prometheus
CloudWatch
Azure Monitor
OpenTelemetry
ServiceNow
Jira
Terraform
Ansible
REST APIs
SQL

Education

Bachelor's degree

Tools

AWS
Azure
Kubernetes
Docker
GitHub Actions
Azure DevOps
Jenkins
GitLab
Datadog
Splunk
Dynatrace
Grafana
Prometheus
CloudWatch
Azure Monitor
OpenTelemetry
ServiceNow
Jira
Terraform
Ansible
REST APIs
SQL

Job description

U.S. Bank is seeking an experienced Site Reliability Engineer to lead complex incident resolution, drive reliability initiatives, and collaborate with software, infra, and product teams.

The role emphasizes RCA, monitoring, IaC, CI/CD, and proactive resilience across AWS/Azure/Kubernetes environments. Eligible candidates should have strong leadership and communication skills for cross-functional coordination.

Get your free, confidential resume review.
or drag and drop your file here.