Senior Lead Site Reliability Engineer

JPMorgan Chase & Co.

Jersey City (NJ)

On-site

USD 150,000 - 210,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

JPMorgan Chase & Co. seeks an experienced SRE/DevOps engineer to drive reliability for mission critical systems in a Kubernetes and AWS environment. You will own production health, automate toil, and strengthen security controls across platforms.

You will implement SLIs/SLOs, improve observability, lead incident response, and collaborate with platform teams to scale containerized services and serverless components using Spinnaker, Harness, Terraform.

Qualifications

  • Formal training or certification on software engineering concepts and 5+ years applied experience.
  • Experience in SRE/DevOps/production engineering or equivalent.
  • Hands-on experience operating Kubernetes workloads in production.
  • Practical experience with AWS in production (EKS/ECS/Lambda).
  • Experience with CI/CD and release tooling such as Spinnaker and Harness.
  • Proficiency with IaC (Terraform) and scripting/automation.
  • Strong incident response skills and RCA writing.
  • Solid Linux, networking, and distributed systems troubleshooting.
  • Ability to execute reliability work independently and escalate when needed.

Responsibilities

  • Own production reliability outcomes by managing day-to-day health and risks.
  • Define and evolve SLIs/SLOs and error budgets; tune alerting.
  • Improve observability with metrics, logs, traces, dashboards, and runbooks.
  • Lead incident response and RCAs; ensure corrective actions close.
  • Operate Kubernetes workloads including autoscaling and rollouts.
  • Operate AWS components (EKS, ECS, Lambda) for scaling and resilience.
  • Improve release engineering and deployment safety with Spinnaker/Harness.
  • Build infrastructure as code with Terraform and automation.
  • Strengthen database and data-service reliability across technologies.
  • Embed secure operational practices and data-sensitivity controls.

Skills

SRE/DevOps
Kubernetes
AWS
Terraform
CI/CD
Spinnaker
Harness
Python/Bash/Go
Incident response
Linux
Networking

Tools

Spinnaker
Harness
Terraform
Scripting (Python/Bash/Go)

Job description

There’s nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.

You are an integral part of a team that works to develop high-quality architecture solutions for various software applications and platform products. You drive significant business impact and help shape the target state architecture through your capabilities in multiple architecture domains. You will ensure the platform is reliable, secure, performant, and resilient in production across Kubernetes-based environments and AWS. You will apply SRE principles to drive measurable improvements in availability and latency, reduce operational toil through automation, and strengthen deployment safety and recovery capabilities in close partnership with engineering and platform teams.

Job responsibilities
  • Own production reliability outcomes by managing day-to-day operational health (availability, latency, throughput, error rates), proactively surfacing risks, and driving remediation.
  • Define and evolve service level indicators/service level objectives (SLIs/SLOs) and error budgets; build actionable, customer-impact-aligned alerting and reduce noise through tuning and standardization.
  • Improve end-to-end observability and troubleshooting (metrics, logs, traces), dashboards, and runbooks across Kubernetes and Amazon Web Services (AWS); perform deep technical triage of distributed-system issues.
  • Lead incident response and problem management by participating in on-call, driving triage/mitigation/recovery, completing root cause analyses (RCAs), and ensuring corrective and preventive actions close.
  • Operate Kubernetes workloads including autoscaling, rollout/rollback procedures, resource tuning, and resilience patterns for containerized services.
  • Operate AWS container and serverless components (for example, Amazon Elastic Kubernetes Service/Elastic Container Service/AWS Lambda) with a focus on scaling, retries, and safe failure modes.
  • Improve release engineering and delivery reliability by increasing the safety and repeatability of deployments using Spinnaker and Harness.
  • Build infrastructure as code and environment consistency by developing and maintaining Terraform modules and automation for reliable, repeatable environments.
  • Strengthen database and data-service reliability by partnering with engineering and platform teams to improve reliability patterns across multiple database technologies and data services (for example, DynamoDB, Amazon Simple Storage Service).
  • Embed security and controls into operations by applying secure operational practices and ensuring processes meet required control standards.
  • Lead small-to-medium initiatives end-to-end from proposal through production adoption, using enterprise-authorized AI capabilities to accelerate triage and toil reduction while validating outputs and handling operational data per sensitivity and security requirements.
Required qualifications, capabilities, and skills
  • Formal training or certification on software engineering concepts and 5+ years applied experience
  • Experience in SRE/DevOps/production engineering or equivalent
  • Hands-on experience operating Kubernetes workloads (deployments, scaling, debugging)
  • Practical experience with AWS (EKS, ECS, Lambda, Dynamo DB, S3) in production
  • Experience with CI/CD and release tooling such as Spinnaker and/or Harness
  • Proficiency with Terraform (IaC), and scripting/automation (Python/Bash/Go)
  • Strong incident response skills, RCA writing, and ability to drive remediation work
  • Solid fundamentals in Linux, networking, and troubleshooting distributed systems
  • Ability to independently execute well-scoped reliability work and elevate when needed
  • Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
  • Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
Preferred qualifications, capabilities, and skills
  • Experience implementing SLO programs and alerting aligned to customer journeys
  • Experience with performance testing, capacity planning, and resilience testing (fault injection/chaos, DR exercises)
  • Experience improving operational maturity: standardized runbooks, automated health checks, auto-remediation, and deployment guardrails
  • Experience with fraud screening/decisioning or payment flows
  • Familiarity with database reliability patterns (capacity, backups, failover readiness)
  • Experience with secure operational practices (least privilege, secrets handling)
  • Experience partnering with engineering and platform teams to drive reliability improvements
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veriipro • Charlotte (NC)

On-site
USD 140,000 - 190,000
Staff Site Reliability Engineer, SRE
Staff Site Reliability Engineer, SRE

Jobtailor • California (MO)

On-site
USD 120,000 - 210,000
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) - Internal Kubernetes Container Platform (IKCP)

Koitecc Solutions • Chandler (AZ), Northern (KY)

On-site
USD 140,000 - 200,000
Site Reliability Engineer – Lead
Site Reliability Engineer – Lead

Jobtailor • Arizona

On-site
USD 140,000 - 230,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America • Plano (TX)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Kovoro • Denver (CO), Northern (KY)

Hybrid
USD 150,000 - 190,000
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)
Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Hobbsnews • Chandler (AZ), Northern (KY)

Hybrid
USD 120,000 - 180,000
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

Archer • San Jose (CA)

On-site
USD 160,000 - 210,000