Senior Lead Site Reliability Engineer

JPMorganChase

Jersey City (NJ)

On-site

USD 170,000 - 210,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health benefits
On-site wellness centers
Retirement savings plan
Tuition reimbursement
Mental health support
Financial coaching

Job summary

JPMorganChase is seeking an experienced Senior Lead Site Reliability Engineer to join the Commercial Investment Banking Fraud Prevention team. You will ensure reliability, security, and performance of Kubernetes-based systems on AWS, while driving autonomous improvements in observability, automation, and deployment safety.

The role emphasizes joining a team that shapes target state architecture, reduces toil, and applies SRE principles to scale critical, mission‑critical platforms in production.

Qualifications

  • Formal training or certification on software engineering concepts and 5+ years applied experience.
  • Experience in SRE/DevOps/production engineering or equivalent.
  • Hands-on experience operating Kubernetes workloads (deployments, scaling, debugging).
  • Practical experience with AWS (EKS, ECS, Lambda, Dynamo DB, S3) in production.
  • Experience with CI/CD and release tooling such as Spinnaker and/or Harness.
  • Proficiency with Terraform (IaC), and scripting/automation (Python/Bash/Go).
  • Strong incident response skills, RCA writing, and ability to drive remediation work.
  • Solid fundamentals in Linux, networking, and troubleshooting distributed systems.
  • Ability to independently execute well-scoped reliability work and elevate when needed.

Responsibilities

  • Own production reliability outcomes by managing day-to-day operational health (availability, latency, throughput, error rates).
  • Define and evolve SLIs/SLOs and error budgets; build actionable alerting and reduce noise.
  • Improve observability, dashboards, runbooks across Kubernetes and AWS; triage distributed-system issues.
  • Lead incident response and problem management with on-call participation and RCAs.
  • Operate Kubernetes workloads including autoscaling and rollout/rollback procedures.
  • Operate AWS container and serverless components (EKS/ECS/Lambda) focusing on scaling and safe failure modes.
  • Improve release engineering and delivery reliability with Spinnaker and Harness.
  • Build infrastructure as code with Terraform modules and automation for reliable environments.
  • Strengthen database and data-service reliability with DynamoDB and S3.
  • Embed security and controls into operations; meet required security standards.
  • Lead initiatives end-to-end, leveraging enterprise AI capabilities with proper data sensitivity.

Skills

SRE/DevOps
Kubernetes
AWS
CI/CD
Terraform
Scripting
Incident response

Tools

Spinnaker
Harness
Terraform
Kubernetes

Job description

Job Description

There's nothing more exciting than being at the center of a rapidly growing field in technology and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.

As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Commercial Investment Banking team of Fraud Prevention, you will solve complex and broad business problems with simple and straightforward solutions. Through code and cloud infrastructure, you will configure, maintain, monitor, and optimize applications and their associated infrastructure to independently decompose and iteratively improve on existing solutions. You are a significant contributor to your team by sharing your knowledge of end-to-end operations, availability, reliability, and scalability of your application or platform.

You are an integral part of a team that works to develop high-quality architecture solutions for various software applications and platform products. You drive significant business impact and help shape the target state architecture through your capabilities in multiple architecture domains. You will ensure the platform is reliable, secure, performant, and resilient in production across Kubernetes-based environments and AWS. You will apply SRE principles to drive measurable improvements in availability and latency, reduce operational toil through automation, and strengthen deployment safety and recovery capabilities in close partnership with engineering and platform teams.

Job Responsibilities
  • Own production reliability outcomes by managing day-to-day operational health (availability, latency, throughput, error rates), proactively surfacing risks, and driving remediation.
  • Define and evolve service level indicators/service level objectives (SLIs/SLOs) and error budgets; build actionable, customer-impact-aligned alerting and reduce noise through tuning and standardization.
  • Improve end-to-end observability and troubleshooting (metrics, logs, traces), dashboards, and runbooks across Kubernetes and Amazon Web Services (AWS); perform deep technical triage of distributed-system issues.
  • Lead incident response and problem management by participating in on-call, driving triage/mitigation/recovery, completing root cause analyses (RCAs), and ensuring corrective and preventive actions close.
  • Operate Kubernetes workloads including autoscaling, rollout/rollback procedures, resource tuning, and resilience patterns for containerized services.
  • Operate AWS container and serverless components (for example, Amazon Elastic Kubernetes Service/Elastic Container Service/AWS Lambda) with a focus on scaling, retries, and safe failure modes.
  • Improve release engineering and delivery reliability by increasing the safety and repeatability of deployments using Spinnaker and Harness.
  • Build infrastructure as code and environment consistency by developing and maintaining Terraform modules and automation for reliable, repeatable environments.
  • Strengthen database and data-service reliability by partnering with engineering and platform teams to improve reliability patterns across multiple database technologies and data services (for example, DynamoDB, Amazon Simple Storage Service).
  • Embed security and controls into operations by applying secure operational practices and ensuring processes meet required control standards.
  • Lead small-to-medium initiatives end-to-end from proposal through production adoption, using enterprise-authorized AI capabilities to accelerate triage and toil reduction while validating outputs and handling operational data per sensitivity and security requirements.
Required Qualifications, Capabilities, And Skills
  • Formal training or certification on software engineering concepts and 5+ years applied experience
  • Experience in SRE/DevOps/production engineering or equivalent
  • Hands-on experience operating Kubernetes workloads (deployments, scaling, debugging)
  • Practical experience with AWS (EKS, ECS, Lambda, Dynamo DB, S3) in production
  • Experience with CI/CD and release tooling such as Spinnaker and/or Harness
  • Proficiency with Terraform (IaC), and scripting/automation (Python/Bash/Go)
  • Strong incident response skills, RCA writing, and ability to drive remediation work
  • Solid fundamentals in Linux, networking, and troubleshooting distributed systems
  • Ability to independently execute well-scoped reliability work and elevate when needed
  • Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
  • Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
Preferred Qualifications, Capabilities, And Skills
  • Experience implementing SLO programs and alerting aligned to customer journeys
  • Experience with performance testing, capacity planning, and resilience testing (fault injection/chaos, DR exercises)
  • Experience improving operational maturity: standardized runbooks, automated health checks, auto-remediation, and deployment guardrails
  • Experience with fraud screening/decisioning or payment flows
  • Familiarity with database reliability patterns (capacity, backups, failover readiness)
  • Experience with secure operational practices (least privilege, secrets handling)
  • Experience partnering with engineering and platform teams to drive reliability improvements
ABOUT US

JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world's most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management.

We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.

  • comprehensive health care coverage
  • on-site health and wellness centers
  • a retirement savings plan
  • backup childcare
  • tuition reimbursement
  • mental health support
  • financial coaching
  • more

We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants' and employees' religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.

JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans

About The Team

J.P. Morgan's Commercial & Investment Bank is a global leader across banking, markets, securities services and payments. Corporations, governments and institutions throughout the world entrust us with their business in more than 100 countries. The Commercial & Investment Bank provides strategic advice, raises capital, manages risk and extends liquidity in markets around the world.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

Fairygodboss • Jersey City (NJ)

On-site
USD 150,000 - 230,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • Houston (TX)

On-site
USD 150,000 - 210,000
Health insurance
Retirement plan
Tuition reimbursement
+1
Site Reliability Engineer II
Site Reliability Engineer II

JPMorganChase • Tampa (FL)

On-site
USD 120,000 - 150,000
Comprehensive health care coverage
On-site health and wellness centers
Retirement savings plan
+4
Lead Software Engineer- AWS/ Cloud /Infra
Lead Software Engineer- AWS/ Cloud /Infra

Fairygodboss • Plano (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer II
Site Reliability Engineer II

JPMorganChase • Chicago (IL)

On-site
USD 110,000 - 150,000
Health care coverage
On-site health centers
Retirement savings plan
+4
Lead Site Reliability Engineer
Lead Site Reliability Engineer

J.P. Morgan • New York (NY)

On-site
USD 150,000 - 210,000
Health care coverage
On-site health & wellness centers
Retirement savings plan
+4
Site Reliability Engineer III- Production Management
Site Reliability Engineer III- Production Management

Next Frontier Capital • New York (NY)

On-site
USD 130,000 - 180,000
Health insurance
Retirement plan
On-site wellness centers
Site Reliability Engineer III
Site Reliability Engineer III

JPMorganChase • Plano (TX)

On-site
USD 130,000 - 180,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Next Frontier Capital • Kentucky

On-site
USD 120,000 - 180,000
Health care coverage
On-site wellness centers
Retirement savings plan
+4
Site Reliability Engineer III - AWS, Java and Kubernetes
Site Reliability Engineer III - AWS, Java and Kubernetes

JPMorganChase • Chicago (IL)

On-site
USD 130,000 - 190,000