Senior SRE Lead: Platform Reliability & Kubernetes

Hobbsnews

Chandler, Northern (AZ, KY)

Hybrid

USD 120,000 - 180,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Bank of America seeks an IKCP Site Reliability Engineer Lead to ensure reliability, scalability, performance, security, and operability of the enterprise Internal Kubernetes Container Platform (IKCP). You will drive automation, observability, incident management, capacity planning, and platform governance across OpenShift, Kubernetes, Rancher, and related services.

You will partner with Engineering, Architecture, Product Management, Security, and Central Operations to deliver a highly available

Qualifications

  • 8+ years of infrastructure, cloud, platform engineering, or SRE experience.
  • 5+ years managing Kubernetes and/or OpenShift production environments.
  • Experience operating large-scale mission-critical distributed systems.
  • Experience supporting enterprise production environments with 24x7 operational responsibilities.

Responsibilities

  • Own platform reliability objectives, including service availability, resiliency, recoverability, and operational health.
  • Lead critical incident response, root cause analysis, and problem management activities.
  • Serve as a senior escalation point for L3 platform support and on-call operations.
  • Develop and maintain operational runbooks, recovery procedures, and standard operating practices.
  • Drive production readiness reviews for new platform capabilities and services.
  • Ensure platforms meet enterprise resiliency and availability objectives.
  • Conduct resilience exercises and continuous improvement activities following recovery testing.
  • Execute platform upgrades, patching strategies, cluster modernization, and release orchestration.
  • Improve platform scalability, performance, and resource utilization across environments.
  • Support platform modernization initiatives including OpenShift virtualization, VKS, and cloud-native technologies.
  • Collaborate with Product, Architecture, Engineering, and Operations teams to improve developer experience and platform adoption.
  • Design and implement enterprise observability solutions leveraging monitoring, logging, tracing, and alerting platforms.
  • Automate operational processes using Infrastructure-as-Code, GitOps, CI/CD, and scripting frameworks.

Skills

Automation
Collaboration
Influence
Production Support
Result Orientation
Analytical Thinking
Application Development
Architecture
Solution Design
Stakeholder Management
Adaptability
DevOps Practices
Project Management
Risk Management
Solution Delivery Process

Education

BS /MS degree in Computer Science, Engineering, Information Systems, or related technical discipline, or equivalent experience

Tools

Terraform
Ansible
GitOps
ArgoCD
Helm
Jenkins
GitHub
GitLab
Bitbucket
Dynatrace
Prometheus
Grafana
Splunk
ELK
OpenTelemetry

Job description

Bank of America seeks an IKCP Site Reliability Engineer Lead to ensure reliability, scalability, performance, security, and operability of the enterprise Internal Kubernetes Container Platform (IKCP). You will drive automation, observability, incident management, capacity planning, and platform governance across OpenShift, Kubernetes, Rancher, and related services.

You will partner with Engineering, Architecture, Product Management, Security, and Central Operations to deliver a highly available

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE Lead: Kubernetes & OpenShift Platform
Senior SRE Lead: Kubernetes & OpenShift Platform

Koitecc Solutions • Plano (TX)

On-site
USD 125,000 - 168,000
SRE Lead: Internal Kubernetes Platform & Observability
SRE Lead: Internal Kubernetes Platform & Observability

Bank of America • Plano (TX)

On-site
USD 125,000 - 168,000
Industry-leading benefits
Discretionary incentive plan
SRE Lead: Kubernetes, Observability & Platform Resilience
SRE Lead: Kubernetes, Observability & Platform Resilience

Bank of America • Plano (TX)

On-site
USD 140,000 - 190,000
Lead SRE: Internal Kubernetes Platform & Automation
Lead SRE: Internal Kubernetes Platform & Automation

Koitecc Solutions • Chandler (AZ), Northern (KY)

Hybrid
USD 140,000 - 200,000
Lead SRE: Enterprise Kubernetes Platform
Lead SRE: Enterprise Kubernetes Platform

Bank of America • Jersey City (NJ)

On-site
USD 125,000 - 168,000
Senior SRE Lead: Cloud Platform Reliability & Automation
Senior SRE Lead: Cloud Platform Reliability & Automation

Bank of America • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Lead SRE: Platform Reliability for Kubernetes/OpenShift
Lead SRE: Platform Reliability for Kubernetes/OpenShift

Bank of America • Charlotte (NC)

On-site
USD 125,000 - 168,000
Discretionary incentive eligible
Annual discretionary plan
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Hobbsnews • Chandler (AZ), Northern (KY)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Plano (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Jersey City (NJ)

On-site
USD 180,000 - 240,000