Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Bank of America

United States

On-site

USD 140,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Bank of America is seeking an IKCP Site Reliability Engineer Lead to ensure the reliability, scalability, and security of the enterprise Kubernetes platform. You will lead automation, observability, incident management, capacity planning, and governance across OpenShift, Kubernetes, Rancher, and VKS.

You will partner with Engineering, Architecture, Security, and Operations to deliver a platform‑as‑a‑product experience for application teams, drive SLOs and runbooks, and mentor junior engineers in

Qualifications

  • 8+ years of infrastructure, cloud, platform engineering, or SRE experience.
  • 5+ years managing Kubernetes and/or OpenShift.

Responsibilities

  • Own platform reliability objectives, including service availability, resiliency, recoverability, and operational health.
  • Lead critical incident response, root cause analysis, and problem management activities.
  • Serve as a senior escalation point for L3 platform support and on‑call operations.
  • Develop and maintain operational runbooks, recovery procedures, and standard operating practices.
  • Drive production readiness reviews for new platform capabilities and services.
  • Ensure platforms meet enterprise resiliency and availability objectives.
  • Conduct resilience exercises and continuous improvement activities following recovery testing.

Skills

Kubernetes
OpenShift
Platform engineering
SRE
Incident management
Automation
Observability

Tools

OpenShift
Rancher
VKS

Job description

Job Description:

At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities and shareholders every day.

Being a Great Place to Work and providing a culture of caring is core to how we drive Responsible Growth. We are intentional about fostering an inclusive workplace where every teammate has the opportunity to succeed, build a career and contribute to our shared success. This includes attracting and developing exceptional talent, recognizing and rewarding performance, and supporting our teammates' physical, emotional, and financial wellness through affordable, competitive and flexible benefits.

We value the unique perspectives individuals bring from all backgrounds and career paths - whether shaped by military service, community college education, or a wide range of work and life experiences. These journeys foster resilience, leadership and innovation, strengthening our workforce and positively impact the communities we serve.

Bank of America is committed to an in-office culture that supports collaboration, engagement, and career development. Our approach includes clear in-office expectations, while providing an appropriate level of flexibility based on role‑specific responsibilities and business needs.

At Bank of America, you can build a successful career with opportunities to learn, grow, and make an impact. Join us!

Position Summary:

The IKCP Site Reliability Engineer Lead is responsible for ensuring the reliability, scalability, performance, security, and operational excellence of the enterprise Internal Kubernetes Container Platform (IKCP). This role serves as a technical lead within the platform organization, driving automation, observability, incident management, capacity planning, platform resilience, and continuous improvement across OpenShift, Kubernetes, Rancher, VKS and emerging container platform services.

The role partners closely with Engineering, Architecture, Product Management, Security, Infrastructure, and Central Operations teams to deliver a highly available platform‑as‑a‑product experience for application teams. Responsibilities are aligned with IKCP's focus on SLOs, error budgets, observability, runbooks, L3 operations, upgrade orchestration, and platform governance.

Key Responsibilities:
Reliability & Operations
  • Own platform reliability objectives, including service availability, resiliency, recoverability, and operational health.
  • Lead critical incident response, root cause analysis, and problem management activities.
  • Serve as a senior escalation point for L3 platform support and on‑call operations.
  • Develop and maintain operational runbooks, recovery procedures, and standard operating practices.
  • Drive production readiness reviews for new platform capabilities and services.
  • Ensure platforms meet enterprise resiliency and availability objectives.
  • Conduct resilience exercises and continuous improvement activities following recovery testing.
Kubernetes & OpenShift Platform Engineering
  • Execute platform upgrades, patching strategies, cluster modernization, and release orchestration.
  • Improve platform scalability, performance, and resource utilization across production and non‑production environments.
  • Support platform modernization initiatives including OpenShift virtualization, VKS, and cloud‑native technologies
  • Collaborate with Product, Architecture, Engineering, and Operations teams to improve developer experience and platform adoption.
Observability & Automation
  • Design and implement enterprise observability solutions leveraging monitoring, logging, tracing, and alerting platforms.
  • Automate operational processes using Infrastructure‑as‑Code, GitOps, CI/CD, and scripting frameworks.
  • Reduce operational toil through self‑healing, intelligent automation, and proactive remediation capabilities.
  • Drive operational efficiency through automation of cluster provisioning, upgrades, compliance, and day-2 operations.
Capacity & Performance Engineering
  • Perform platform capacity planning and trend analysis.
  • Forecast infrastructure growth requirements and optimize platform resource consumption.
  • Conduct performance tuning for clusters, workloads, networking, and storage services.
  • Support enterprise‑scale growth while maintaining platform stability and customer experience.
Security & Compliance
  • Partner with security teams to implement platform security controls and governance requirements.
  • Support vulnerability remediation, image compliance, platform hardening, and policy enforcement.
  • Implement and maintain RBAC, Network Policies, and container security controls.
  • Drive compliance with enterprise standards, vulnerability management processes, and audit requirements.
Required Qualifications:
Education / Experience
  • 8+ years of infrastructure, cloud, platform engineering, or SRE experience.
  • 5+ years managing Kubernetes and/or OpenShif
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Hobbsnews • Chandler (AZ), Northern (KY)

Hybrid
USD 120,000 - 180,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Plano (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Koitecc Solutions • Chandler (AZ), Northern (KY)

On-site
USD 140,000 - 200,000
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Bank of America • Jersey City (NJ)

On-site
USD 125,000 - 168,000
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Koitecc Solutions • Plano (TX)

On-site
USD 125,000 - 168,000
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Bank of America • Plano (TX)

On-site
USD 125,000 - 168,000
Industry-leading benefits
Discretionary incentive plan
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Charlotte (NC)

On-site
USD 125,000 - 168,000
Discretionary incentive eligible
Annual discretionary plan
Senior SRE Lead: Platform Reliability & Kubernetes
Senior SRE Lead: Platform Reliability & Kubernetes

Hobbsnews • Chandler (AZ), Northern (KY)

Hybrid
USD 120,000 - 180,000
SRE Lead: Internal Kubernetes Platform & Observability
SRE Lead: Internal Kubernetes Platform & Observability

Bank of America • Plano (TX)

On-site
USD 125,000 - 168,000
Industry-leading benefits
Discretionary incentive plan