Lead Site Reliability Engineer

Next Frontier Capital

Kentucky

On-site

USD 120,000 - 180,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health care coverage
On-site wellness centers
Retirement savings plan
Backup childcare
Tuition reimbursement
Mental health support
Financial coaching

Job summary

JPMorgan Chase & Co. seeks an experienced Site Reliability Engineer to define the future of reliability for a globally recognized firm. You will guide peers, design and implement CI/CD pipelines, and build scalable, observable systems with strong AI-enabled triage capabilities.

The role emphasizes hands-on engineering across containers, cloud, and networking with 24/7 production support and a focus on security, resiliency, and auditability in the SDLC.

Qualifications

  • Formal training or certification on site reliability engineering concepts and 5+ years applied experience.
  • Proficient in SRE culture and principles, with strong observability including SLO-based alerting and telemetry collection.
  • Proficient in at least one programming language with knowledge of cloud, AI, or related areas.
  • Demonstrated experience using enterprise AI capabilities to improve SRE workflows with validation habits.

Responsibilities

  • Lead and guide peers in designing scalable level designs and achieving consensus.
  • Collaborate to design and implement CI/CD pipelines for reliable delivery.
  • Build and improve availability, reliability, and scalability across applications.
  • Implement infrastructure as code and network automation for platforms.
  • Collaborate with stakeholders to resolve complex issues.
  • Monitor service level indicators and use SLOs to prevent customer impact.
  • Advocate for SRE best practices and practical AI-assisted workflows.
  • Provide 24/7 production support for business-critical apps.
  • Leverage AI to accelerate major-incident triage and post-incident analysis.
  • Promote traceability, resiliency, and security controls across toolchains.

Skills

SRE principles
Observability
Programming: Python/Java/.NET
AI for SRE
CI/CD tooling
Containers/Kubernetes
Networking troubleshooting
Team collaboration
Kafka experience

Education

Formal SRE training/certification

Tools

Grafana
Dynatrace
Prometheus
Datadog
Splunk
Jenkins
GitLab
Terraform
Docker
Kubernetes
ECS

Job description

Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.

Job responsibilities
  • Guides and assists others in the areas of building appropriate level designs and gaining consensus from peers where appropriate
  • Collaborates with other software engineers and teams to design and implement deployment approaches using automated continuous integration and continuous delivery pipelines
  • Collaborates with other software engineers and teams to design, develop, test, and implement availability, reliability, scalability, and solutions in their applications
  • Implements infrastructure, configuration, and network as code for the applications and platforms in your remit
  • Collaborates with technical experts, key stakeholders, and team members to resolve complex problems
  • Understands service level indicators and utilizes service level objectives to proactively resolve issues before they impact customers
  • Supports the adoption of site reliability engineering best practices within your team
  • Production 24*7 support for business-critical applications
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.
Required qualifications, capabilities, and skills
  • Formal training or certification on site reliability engineering concepts and 5+ years applied experience
  • Proficient in site reliability engineering (SRE) culture and principles, with experience implementing SRE practices within applications and platforms; strong observability background including white/black-box monitoring, SLO-based alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, and similar.
  • Proficient in at least one programming language (e.g., Python, Java/Spring Boot,.NET) with strong knowledge of software applications and technical processes within a technical discipline such as cloud, artificial intelligence, Android, or related areas.
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
  • Hands-on experience with CI/CD tooling (e.g., Jenkins, GitLab) and infrastructure automation using Terraform to build reliable, repeatable delivery pipelines.
  • Strong familiarity with containers and orchestration platforms (Docker, Kubernetes, ECS), including deploying, scaling, and operating containerized services in production.
  • Proven ability to troubleshoot and resolve common networking issues (DNS, TCP/IP, routing, TLS, load balancing), applying structured debugging to restore service quickly.
  • Collaborative, proactive team contributor: communicates clearly and persuasively with minimal supervision, identifies roadblocks early, learns new technologies quickly, and has experience with event streaming platforms such as Kafka.
Preferred qualifications, capabilities, and skills
  • Ability to identify new technologies and relevant solutions to ensure design constraints are met by the software team
  • Proven track record of initiating and executing ideas that address complex business challenges
  • Deep expertise in networking and systems, including TCP/IP, DNS, load balancing, firewalls, and VPN technologies; strong Linux performance tuning and system-level troubleshooting skills
  • Certifications a plus: AWS Certified SysOps Administrator or AWS Professional, Certified Kubernetes Administrator (CKA), Terraform Associate (or equivalent)
  • Collaborative leader with a proven track record mentoring junior engineers, driving SRE best-practice adoption across teams, and communicating clearly to both technical and non-technical stakeholders (including presentations)
  • Experience in handling critical incident and change management – be part of critical incident taskforce call.
  • Familiarity of agile practices – preferably, scrum and Kanban

JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world’s most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management.

We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility.

  • comprehensive health care coverage
  • on-site health and wellness centers
  • a retirement savings plan
  • backup childcare
  • tuition reimbursement
  • mental health support
  • financial coaching and more

We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.

JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans

Our professionals in our Corporate Functions cover a diverse range of areas from finance and risk to human resources and marketing. Our corporate teams are an essential part of our company, ensuring that we’re setting our businesses, clients, customers and employees up for success. Lead and conduct resiliency design reviews, break up complex problems, and act as a technical lead for medium to large sized products

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • Jersey City (NJ)

On-site
USD 170,000 - 250,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • Jersey City (NJ)

On-site
USD 170,000 - 230,000
AWS Certifications
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • Houston (TX)

On-site
USD 150,000 - 210,000
Health insurance
Retirement plan
Tuition reimbursement
+1
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

Fairygodboss • Plano (TX)

On-site
USD 150,000 - 190,000
Comprehensive health care coverage
On-site health and wellness centers
Retirement savings plan
+4
Lead Site Reliability Engineer - Operations Excellence for AI Platforms
Lead Site Reliability Engineer - Operations Excellence for AI Platforms

JPMorganChase • Jersey City (NJ)

On-site
USD 140,000 - 200,000
Lead Site Reliability Engineer - Network
Lead Site Reliability Engineer - Network

Fairygodboss • Columbus (OH)

On-site
USD 150,000 - 210,000
Lead Site Reliability Engineer - Network
Lead Site Reliability Engineer - Network

JPMorganChase • Columbus (OH)

On-site
USD 150,000 - 230,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

J.P. Morgan • New York (NY)

On-site
USD 150,000 - 210,000
Health care coverage
On-site health & wellness centers
Retirement savings plan
+4
Site Reliability Engineer III - Machine Learning
Site Reliability Engineer III - Machine Learning

JPMorganChase • Wilmington (DE)

On-site
USD 120,000 - 155,000
Senior Lead Site Reliability Engineer - AI/ML and Data Platforms
Senior Lead Site Reliability Engineer - AI/ML and Data Platforms

Next Frontier Capital • Jersey City (NJ)

On-site
USD 180,000 - 240,000