SRE Lead: Incident Commander for Cloud Reliability

Us Bank

Atlanta (GA)

On-site

USD 112,000 - 131,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Healthcare
Life Insurance
Disability
Parental leave
401(k)
Paid vacation
Holidays
Adoption assistance
Disability leave

Job summary

U.S. Bank is seeking an experienced Site Reliability Engineer/DevOps specialist to lead the resolution of complex production incidents and drive reliability across cloud platforms.

You will mentor teams, coordinate incident response as Incident Commander, and partner with software engineering, infrastructure, and product teams to remediate recurring issues. Applicants should have 6–8 years of relevant experience, a Bachelor's degree or equivalent, and deep knowledge of AWS/Azure, Kubernetes,

Qualifications

  • Bachelor's degree or equivalent work experience.
  • Six to eight years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development.

Responsibilities

  • Lead troubleshooting and resolution of complex production incidents, including application failures, API issues, cloud outages, performance degradation, and disruptions.
  • Conduct comprehensive root cause analysis (RCA), impact assessments, mitigation planning, and implementation of permanent corrective actions.
  • Design and enhance monitoring, observability, alerting, dashboards, health checks, and runbooks to improve platform reliability and availability.
  • Drive automation initiatives using scripting, Infrastructure as Code (IaC), CI/CD pipelines, and self-healing capabilities.
  • Partner with software engineering, infrastructure, and product teams to identify, prioritize, and remediate recurring reliability issues.
  • Serve as Incident Commander during major incidents, coordinating cross-functional response teams and driving restoration activities.
  • Provide leadership, coaching, mentoring, and workload management for SRE, DevOps, and production support engineers.
  • Utilize MTTR, MTTD, SLA compliance, backlog health, incident volume, and problem closure rates to drive continuous improvement.

Skills

SRE
DevOps
Production Support
Platform Engineering
Distributed Systems
Incident Management
Change Management
RCA
AWS
Azure
Kubernetes
Docker
Python
CI/CD
Observability
ServiceNow
Jira
Terraform
Ansible
REST APIs
SQL
Datadog
Splunk
Dynatrace
Grafana
Prometheus
CloudWatch
Azure Monitor
OpenTelemetry

Education

Bachelor's degree, or equivalent

Tools

Kubernetes
Docker
AWS
Azure
GitHub Actions
Azure DevOps
Jenkins
GitLab
Terraform
Ansible
REST APIs
SQL/Relational Databases
Datadog
Splunk
Dynatrace
Grafana
Prometheus
CloudWatch
Azure Monitor
OpenTelemetry
ServiceNow
Jira

Job description

U.S. Bank is seeking an experienced Site Reliability Engineer/DevOps specialist to lead the resolution of complex production incidents and drive reliability across cloud platforms.

You will mentor teams, coordinate incident response as Incident Commander, and partner with software engineering, infrastructure, and product teams to remediate recurring issues. Applicants should have 6–8 years of relevant experience, a Bachelor's degree or equivalent, and deep knowledge of AWS/Azure, Kubernetes,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Lead - Incident Commander & Reliability Architect
SRE Lead - Incident Commander & Reliability Architect

U.S. Bank • Atlanta (GA)

On-site
USD 112,000 - 131,000
Healthcare
401(k)
Paid vacation
+1
SRE Lead: Incident Commander & Reliability Champion
SRE Lead: Incident Commander & Reliability Champion

U.S. Bank • Northern (KY)

Hybrid
USD 112,000 - 131,000
Healthcare
Retirement plan
Paid vacation
+2
SRE Lead: Reliability, Incident Command & Automation
SRE Lead: Reliability, Incident Command & Automation

Relha LLC • Atlanta (GA), Northern (KY)

Hybrid
USD 112,000 - 131,000
Life insurance
Disability
Parental leave
+5
Senior SRE Lead: Cloud Platform Reliability & Automation
Senior SRE Lead: Cloud Platform Reliability & Automation

Bank of America • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Lead SRE: Scale, Resilience & Cloud Automation
Lead SRE: Scale, Resilience & Cloud Automation

Federal Reserve Bank of New York • San Francisco (CA)

On-site
USD 147,000 - 234,000
Senior Cloud SRE: AWS, Serverless & Incident Response
Senior Cloud SRE: AWS, Serverless & Incident Response

Apply • Northern (KY)

Hybrid
USD 120,000 - 150,000
Lead SRE: Cloud Reliability & DevOps Leadership
Lead SRE: Cloud Reliability & DevOps Leadership

Federal Reserve Bank of San Francisco • Richmond (VA)

On-site
USD 147,000 - 234,000
SRE Engineering Manager: Lead Reliability & Incident Response
SRE Engineering Manager: Lead Reliability & Incident Response

PNC • Phoenix (AZ)

On-site
USD 150,000 - 210,000
Medical insurance
Prescription drug coverage
Health Savings Account
+7
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

myBridge Corporation • Austin (TX)

On-site
USD 120,000 - 160,000
SRE Lead: Reliability, AI-Driven Incident Mastery
SRE Lead: Reliability, AI-Driven Incident Mastery

JPMorganChase • Jersey City (NJ)

On-site
USD 170,000 - 250,000