Lead Site Reliability Engineer

JPMorgan Chase & Co.

Wilmington (DE)

On-site

USD 140,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPMorgan Chase & Co. is seeking a Lead Site Reliability Engineer in Wilmington, DE, to drive resiliency across critical software platforms. You will mentor engineers, lead incident response, and shape target-state architectures with expertise in multiple domains.

You will define SLIs/SLOs, manage major incidents, and champion AI-enabled reliability workflows while ensuring security and governance in a highly regulated banking environment. This role offers significant impact and growth.

Qualifications

  • Formal training or certification in software engineering.
  • 5+ years as an SRE and 10+ years in a regulated industry such as Banking.
  • Experience designing, deploying, and supporting highly available services in a public cloud environment (AWS, Azure, or GCP).
  • Proficiency in observability tools (Grafana, Dynatrace, Prometheus, Datadog, Splunk).
  • Experience with CI/CD tools (Jenkins, GitLab, Terraform).
  • Experience with containers and orchestration (Docker, Kubernetes, ECS).
  • Demonstrated experience using enterprise-authorized AI capabilities within work to improve SRE workflows.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk.

Responsibilities

  • Lead resiliency design reviews and mentor engineers.
  • Define service level indicators and work with stakeholders to establish service level objectives and error budgets.
  • Act as main incident contact during major incidents and drive rapid resolution.
  • Provide technical leadership for medium to large-sized products.
  • Lead adoption of AI-assisted reliability workflows across SDLC/toolchains.
  • Document and share knowledge within the organization.
  • Collaborate with stakeholder partners to establish SLIs and error budgets.

Skills

Site reliability engineering
Leadership
Cloud platforms (AWS/Azure/GCP)
Observability tools
CI/CD tools
Containers & orchestration
Incident management
AI in SRE governance

Education

Software engineering certification

Tools

AWS
Azure
GCP

Job description

Overview

As a Lead Site Reliability Engineer at JPMorganChase within the Corporate sector, Enterprise Technology team, you are an integral part of a team that develops high-quality architecture solutions for critical software applications and platforms. You will lead resiliency design reviews, break down complex problems, and mentor engineers, driving significant business impact and shaping the target state architecture through your expertise in multiple architecture domains. Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.

Responsibilities
  • Demonstrate and champion site reliability culture and practices, exerting technical influence across your team
  • Lead initiatives to improve reliability and stability of applications and platforms using data-driven analytics
  • Collaborate with team members to define service level indicators and work with stakeholders to establish service level objectives and error budgets
  • Provide technical leadership and guidance for medium to large-sized products
  • Proactively identify and resolve technology-related bottlenecks in your areas of expertise
  • Act as the main point of contact during major incidents, quickly identifying and solving issues to avoid financial losses
  • Document and share knowledge within the organization through internal forums and communities of practice
  • Use enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements
  • Lead reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls
  • Offer mentorship and advice to other engineers, fostering a culture of continuous improvement
  • Drive collaboration with stakeholder partners to establish reasonable service level objectives and error budgets
Required qualifications, capabilities, and skills
  • Formal training or certification on software engineering concepts and 5+ years applied experience
  • At least 5 years as an SRE and at least 10 years in a highly regulated industry such as Banking
  • Deep proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and site reliability best practices, with the ability to implement these practices within an application or platform
  • Demonstrated experience designing, deploying, and supporting highly available services in a public cloud environment (AWS, Azure, or GCP); familiarity with cloud-native observability, auto-scaling, and infrastructure-as-code is essential
  • Fluency in at least one programming language (e.g., Python, Java Spring Boot, .Net)
  • Deep knowledge of software applications and technical processes with emerging depth in one or more technical disciplines
  • Proficiency and experience in observability, including white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk
  • Proficiency in continuous integration and continuous delivery tools (e.g., Jenkins, GitLab, Terraform)
  • Experience with containers and container orchestration (e.g., ECS, Kubernetes, Docker)
  • Experience troubleshooting common networking technologies and issues
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations
Preferred qualifications, capabilities, and skills
  • Ability to identify and solve problems related to complex data structures and algorithms
  • Drive to self-educate and evaluate new technology
  • Ability to teach new programming languages to team members
  • Ability to expand and collaborate across different levels and stakeholder groups
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorgan Chase & Co. • Kentucky

On-site
USD 150,000 - 210,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorgan Chase & Co. • Columbus (OH)

On-site
USD 150,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • Plano (TX)

On-site
USD 150,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

慨正橡扯 • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Lead Site Reliability Engineer Market Risk
Lead Site Reliability Engineer Market Risk

JPMorgan Chase & Co. • Houston (TX)

On-site
USD 120,000 - 150,000
Site Reliability Engineer III - Machine Learning
Site Reliability Engineer III - Machine Learning

JPMorgan Chase & Co. • Wilmington (DE)

On-site
USD 120,000 - 160,000
Site Reliability Engineer III
Site Reliability Engineer III

JPMorgan Chase & Co. • Houston (TX)

On-site
USD 120,000 - 180,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • Columbus (OH)

On-site
USD 150,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • Jersey City (NJ)

On-site
USD 170,000 - 230,000
AWS Certifications
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorgan Chase & Co. • City of Rochester (NY)

On-site
USD 170,000 - 230,000