Site Reliability Engineering (SRE ) lead

U.S. Bank

Chicago (IL)

On-site

USD 112,000 - 131,000

Full time

4 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
401(k) retirement plan
Paid vacation

Job summary

U.S. Bank is seeking an experienced Senior Site Reliability Engineer to lead incident response, RCA, and reliability programs. You will drive monitoring, automation, and IaC across AWS/Azure, Kubernetes, and CI/CD stacks while mentoring DevOps and production support engineers.

The role emphasizes cross-functional collaboration with software and infrastructure teams, strong governance, and data-driven improvement of system reliability at scale. Eligible location is a U.S.

Qualifications

  • Bachelor's degree or equivalent work experience.
  • Six to eight years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development.

Responsibilities

  • Lead the troubleshooting and resolution of complex production incidents, including application failures, API issues, cloud platform outages, performance degradation, and operational disruptions.
  • Conduct comprehensive root cause analysis (RCA), impact assessments, mitigation planning, and implementation of permanent corrective actions.
  • Design and enhance monitoring, observability, alerting, dashboards, health checks, and operational runbooks to improve platform reliability and availability.
  • Drive automation initiatives using scripting, Infrastructure as Code (IaC), CI/CD pipelines, and self-healing capabilities to reduce manual operational effort.
  • Partner with software engineering, infrastructure, and product teams to identify, prioritize, and remediate recurring reliability issues.
  • Serve as the Incident Commander during major incidents, coordinating cross-functional response teams and driving restoration activities.
  • Provide leadership, coaching, mentoring, and workload management for SRE, DevOps, and production support engineers.
  • Utilize operational metrics including MTTR, MTTD, SLA compliance, backlog health, incident volume, and problem closure rates to drive continuous improvement and operational excellence.

Skills

SRE
DevOps
Production Support
Platform Engineering
Distributed Systems
Incident response
Workload prioritization
Reliability improvement
Incident Management
Problem Management
Change Management
RCA
Python
PowerShell
Shell Scripting
Automation frameworks
CI/CD tools
GitHub Actions
Azure DevOps
Jenkins
GitLab
Datadog
Splunk
Dynatrace
Grafana
Prometheus
CloudWatch
Azure Monitor
OpenTelemetry
ServiceNow
Jira
Terraform
Ansible
REST APIs
SQL/Relational DBs

Education

Bachelor's degree or equivalent

Tools

AWS
Azure
Kubernetes
Docker

Job description

At U.S. Bank, we’re on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it takes all of us to bring our shared ambition to life, and each person is unique in their potential. A career with U.S. Bank gives you a wide, ever-growing range of opportunities to discover what makes you thrive at every stage of your career. Try new things, learn new skills and discover what you excel at—all from Day One.

Responsibilities
Job Description
  • Lead the troubleshooting and resolution of complex production incidents, including application failures, API issues, cloud platform outages, performance degradation, and operational disruptions.
  • Conduct comprehensive root cause analysis (RCA), impact assessments, mitigation planning, and implementation of permanent corrective actions.
  • Design and enhance monitoring, observability, alerting, dashboards, health checks, and operational runbooks to improve platform reliability and availability.
  • Drive automation initiatives using scripting, Infrastructure as Code (IaC), CI/CD pipelines, and self-healing capabilities to reduce manual operational effort.
  • Partner with software engineering, infrastructure, and product teams to identify, prioritize, and remediate recurring reliability issues.
  • Serve as the Incident Commander during major incidents, coordinating cross-functional response teams and driving restoration activities.
  • Provide leadership, coaching, mentoring, and workload management for SRE, DevOps, and production support engineers.
  • Utilize operational metrics including MTTR, MTTD, SLA compliance, backlog health, incident volume, and problem closure rates to drive continuous improvement and operational excellence.
Basic Qualifications
  • Bachelor's degree, or equivalent work experience
  • Six to eight years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development
Preferred Skills/Experience
  • Strong expertise in Site Reliability Engineering (SRE), DevOps, Production Support, Platform Engineering, and Distributed Systems Operations.
  • Experience leading technical teams, incident response efforts, workload prioritization, and reliability improvement programs.
  • Advanced knowledge of Incident Management, Problem Management, Change Management, and Root Cause Analysis (RCA) methodologies.
  • Hands-on experience with AWS, Azure, Kubernetes, Docker, and cloud-native infrastructure platforms.
  • Proficiency with Python, PowerShell, Shell Scripting, and automation frameworks for operational efficiency and reliability engineering.
  • Experience building and supporting CI/CD pipelines using tools such as GitHub Actions, Azure DevOps, Jenkins, or GitLab.
  • Strong expertise in Monitoring and Observability Solutions including Datadog, Splunk, Dynatrace, Grafana, Prometheus, CloudWatch, Azure Monitor, and OpenTelemetry.
  • Experience with ServiceNow, Jira, Terraform, Ansible, REST APIs, SQL/Relational Databases, along with excellent stakeholder communication and leadership skills.
Preferred Certifications
  • AWS Certified Solutions Architect, DevOps Engineer, or equivalent AWS certification
  • Microsoft Azure Administrator, Architect, or DevOps Engineer certification
  • Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD)
Location expectations

This role requires working from a U.S. Bank location three (3) or more days per week.

If there’s anything we can do to accommodate a disability during any portion of the application or hiring process, please refer to our disability accommodations for applicants.

Benefits:

Our approach to benefits and total rewards considers our team members’ whole selves and what may be needed to thrive in and outside work. That's why our benefits are designed to help you and your family boost your health, protect your financial security and give you peace of mind. Our benefits include the following:

  • Healthcare (medical, dental, vision)
  • Basic term and optional term life insurance
  • Short-term and long-term disability
  • Pregnancy disability and parental leave
  • 401(k) and employer-funded retirement plan
  • Paid vacation (from two to five weeks depending on salary grade and tenure)
  • Up to 11 paid holiday opportunities
  • Adoption assistance
  • Sick and Safe Leave accruals of one hour for every 30 worked, up to 80 hours per calendar year unless otherwise provided by law

Review our full benefits available by employment status here.

E-Verify

U.S. Bank participates in the U.S. Department of Homeland Security E-Verify program in all facilities located in the United States and certain U.S. territories. The E-Verify program is an Internet-based employment eligibility verification system operated by the U.S. Citizenship and Immigration Services. Learn more about the E-Verify program.

The salary range reflects figures based on the primary location, which is listed first. The actual range for the role may differ based on the location of the role. In addition to salary, U.S. Bank offers a comprehensive benefits package, including incentive and recognition programs, equity stock purchase 401(k) contribution and pension (all benefits are subject to eligibility requirements). Pay Range: $111,605.00 - $131,300.00

U.S. Bank will consider qualified applicants with arrest or conviction records for employment. U.S. Bank conducts background checks consistent with applicable local laws, including the Los Angeles County Fair Chance Ordinance and the California Fair Chance Act as well as the San Francisco Fair Chance Ordinance. U.S. Bank is subject to, and conducts background checks consistent with the requirements of Section 19 of the Federal Deposit Insurance Act (FDIA). In addition, certain positions may also be subject to the requirements of FINRA, NMLS registration, Reg Z, Reg G, OFAC, the NFA, the FCPA, the Bank Secrecy Act, the SAFE Act, and/or federal guidelines applicable to an agreement, such as those related to ethics, safety, or operational procedures.

Applicants must be able to comply with U.S. Bank policies and procedures including the Code of Ethics and Business Conduct and related workplace conduct and safety policies.

U.S. Bank is an equal opportunity employer. We consider all qualified applicants without regard to race, religion, color, sex, national origin, age, sexual orientation, gender identity, disability or veteran status, and other factors protected under applicable law.

Posting may be closed earlier due to high volume of applicants.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineering (SRE ) lead
Site Reliability Engineering (SRE ) lead

Us Bank • Atlanta (GA)

On-site
USD 112,000 - 131,000
Healthcare
Life Insurance
Disability
+6
Site Reliability Engineering (SRE ) lead
Site Reliability Engineering (SRE ) lead

Relha LLC • Atlanta (GA), Northern (KY)

Hybrid
USD 112,000 - 131,000
Life insurance
Disability
Parental leave
+5
Site Reliability Engineering (SRE ) lead
Site Reliability Engineering (SRE ) lead

U.S. Bank • Northern (KY)

Hybrid
USD 112,000 - 131,000
Healthcare
Retirement plan
Paid vacation
+2
Site Reliability Engineering (SRE ) lead
Site Reliability Engineering (SRE ) lead

U.S. Bank • Atlanta (GA)

On-site
USD 112,000 - 131,000
Healthcare
401(k)
Paid vacation
+1
Senior Systems Engineer
Senior Systems Engineer

U.S. Bank • Hopkins (MN)

On-site
USD 119,000 - 141,000
Healthcare (medical, dental, vision)
401(k) and employer-funded retirement plan
Paid vacation and holidays
Reliability Engineer 3 (Observability Specialist)
Reliability Engineer 3 (Observability Specialist)

Relha LLC • Town of Brookfield (WI)

Hybrid
USD 98,000 - 116,000
Reliability Engineer 3 (Observability Specialist)
Reliability Engineer 3 (Observability Specialist)

Us Bank • Cincinnati (OH)

On-site
USD 98,000 - 116,000
Healthcare
401(k) plan
Paid vacation
+1
Reliability Engineer 3 (Observability Specialist)
Reliability Engineer 3 (Observability Specialist)

Us Bank • Town of Brookfield (WI)

On-site
USD 98,000 - 116,000
Healthcare
401(k)
Paid vacation
+3
Reliability Engineer 3 (Observability Specialist)
Reliability Engineer 3 (Observability Specialist)

U.S. Bank • Cupertino (CA)

On-site
USD 98,000 - 116,000
Healthcare
Retirement plan
Paid time off
+1
Software Engineer 1 (Java, Linux, SQL)
Software Engineer 1 (Java, Linux, SQL)

Us Bank • Saint Paul (MN)

On-site
USD 93,000 - 109,000
Healthcare
Life insurance
Disability insurance
+6