Lead Site Reliability Engineer

JPMorgan Chase

Kentucky

On-site

USD 150,000 - 230,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

JPMorgan Chase seeks a Lead Site Reliability Engineer to define the future of its enterprise platform. You will guide resiliency reviews, break complex problems into actionable work, and mentor engineers across medium to large products.

You will lead availability and performance design, implement infrastructure as code, and drive SRE best practices with a focus on security and reliability in production.

Qualifications

  • Formal training or certification in site reliability engineering.
  • Strong observability skills with SRE practices and telemetry tools.
  • Proficient in at least one programming language (Python/Java/.NET).
  • Experience using enterprise AI capabilities to improve SRE workflows.
  • Ability to evaluate AI recommendations for correctness and risk.
  • Hands-on experience with CI/CD tooling and Terraform.
  • Familiarity with containers and orchestration (Docker/Kubernetes/ECS).
  • Troubleshoot networking issues (DNS/TCP/IP/TLS/load balancing).
  • Collaborative communicator with cross-team coordination and Kafka experience.

Responsibilities

  • Guides and assists others in designing level designs and consensus.
  • Collaborates to design and implement deployment approaches using CI/CD pipelines.
  • Designs, develops, tests, and implements availability, reliability, scalability solutions.
  • Implements infrastructure, configuration, and network as code.
  • Resolves complex problems with stakeholders and peers.
  • Understands and uses SLI/SLO metrics to preempt issues.
  • Supports adoption of SRE best practices within the team.
  • Provides 24x7 production support for business-critical apps.
  • Leads AI-assisted reliability workflows across SDLC and tooling.
  • Ensures traceability, auditability, resiliency, and security controls.

Skills

SRE concepts
Observability
Programming
AI integration
CI/CD
Terraform
Docker/Kubernetes
Networking
Kafka
Team collaboration

Tools

Grafana
Dynatrace
Prometheus
Datadog
Splunk
Jenkins
GitLab

Job description

Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.

As a Lead Site Reliability Engineer at JPMorgan Chase within the Enterprise technology, engineering services and platform team, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them. Take lead and conduct resiliency design reviews, break up complex problems into digestible work for other engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to other engineers.

Job responsibilities
  • Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate
  • Collaborates with other software engineers and teams to design and implement deployment approaches using automated continuous integration and continuous delivery pipelines
  • Collaborates with other software engineers and teams to design, develop, test, and implement availability, reliability, scalability, and solutions in their applications
  • Implements infrastructure, configuration, and network as code for the applications and platforms in your remit
  • Collaborates with technical experts, key stakeholders, and team members to resolve complex problems
  • Understands service level indicators and utilizes service level objectives to proactively resolve issues before they impact customers
  • Supports the adoption of site reliability engineering best practices within your team
  • Production 24*7 support for business-critical applications
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.
Required qualifications, capabilities, and skills
  • Formal training or certification on site reliability engineering concepts and 5+ years applied experience
  • Proficient in site reliability engineering (SRE) culture and principles, with experience implementing SRE practices within applications and platforms; strong observability background including white/black-box monitoring, SLO-based alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, and similar.
  • Proficient in at least one programming language (e.g., Python, Java/Spring Boot,.NET) with strong knowledge of software applications and technical processes within a technical discipline such as cloud, artificial intelligence, Android, or related areas.
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
  • Hands-on experience with CI/CD tooling (e.g., Jenkins, GitLab) and infrastructure automation using Terraform to build reliable, repeatable delivery pipelines.
  • Strong familiarity with containers and orchestration platforms (Docker, Kubernetes, ECS), including deploying, scaling, and operating containerized services in production.
  • Proven ability to troubleshoot and resolve common networking issues (DNS, TCP/IP, routing, TLS, load balancing), applying structured debugging to restore service quickly.
  • Collaborative, proactive team contributor: communicates clearly and persuasively with minimal supervision, identifies roadblocks early, learns new technologies quickly, and has experience with event streaming platforms such as Kafka.
Preferred qualifications, capabilities, and skills
  • Ability to identify new technologies and relevant solutions to ensure design constraints are met by the software team
  • Proven track record of initiating and executing ideas that address complex business challenges
  • Deep expertise in networking and systems, including TCP/IP, DNS, load balancing, firewalls, and VPN technologies; strong Linux performance tuning and system-level troubleshooting skills
  • Certifications a plus: AWS Certified SysOps Administrator or AWS Professional, Certified Kubernetes Administrator (CKA), Terraform Associate (or equivalent)
  • Collaborative leader with a proven track record mentoring junior engineers, driving SRE best-practice adoption across teams, and communicating clearly to both technical and non-technical stakeholders (including presentations)
  • Experience in handling critic
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorgan Chase & Co. • Kentucky

On-site
USD 150,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorgan Chase & Co. • Plano (TX)

On-site
USD 140,000 - 190,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 140,000 - 210,000
Lead Site Reliability Engineer - Network
Lead Site Reliability Engineer - Network

JPMorgan Chase & Co. • Kentucky

On-site
USD 120,000 - 180,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorgan Chase & Co. • Houston (TX)

On-site
USD 120,000 - 230,000
Lead Site Reliability Engineer - Operations Excellence for AI Platforms
Lead Site Reliability Engineer - Operations Excellence for AI Platforms

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 180,000 - 260,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorgan Chase & Co. • City of Rochester (NY)

On-site
USD 170,000 - 230,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Plano (TX)

On-site
USD 150,000 - 230,000
Site Reliability Engineer III - AWS, Java and Kubernetes
Site Reliability Engineer III - AWS, Java and Kubernetes

JPMorgan Chase & Co. • Chicago (IL)

On-site
USD 120,000 - 180,000
Site Reliability Engineer III
Site Reliability Engineer III

JPMorgan Chase & Co. • Houston (TX)

On-site
USD 120,000 - 180,000