Lead SRE

JPMorganChase

Plano (TX)

On-site

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

JPMorganChase is seeking a Delivery SRE leader to ensure secure, reliable software delivery with strong SDLC discipline and measurable reliability across security applications. You will govern readiness, quality gates, and resilience, partnering with Product Owners and engineering leadership to bake SRE requirements into design and build phases.

You will define and manage SLOs/SLIs, drive post-incident reviews, and lead continuous improvement across distributed systems, cloud platforms (AWS or

Qualifications

  • 5+ years supporting critical security-focused applications in large-scale environments and mentoring teams.
  • Experience with monitoring/logging tools (e.g., Splunk, AppDynamics) and dashboard technologies; Splunk Administrator certification desired.
  • Strong grasp of SDLC, secure development, DevOps/CI/CD tooling; capable of implementing top-tier continuous improvement with root-cause analysis and auto-remediation.
  • Effective under pressure; accountable, with excellent stakeholder management and communication skills.
  • Global team collaboration with flexibility to engage during critical incidents outside standard business hours.
  • Experience implementing and managing SLOs/SLIs, error budgets, and operational readiness reviews for distributed systems, including post-incident analysis and resilience improvements.
  • Deep expertise in public cloud platforms (AWS or equivalent), infrastructure automation tools (CloudFormation, Terraform), and capacity planning for large-scale environments; driving DevOps and SRE adoption across teams.

Responsibilities

  • Define and enforce quality gates across requirements, design, secure coding, testing, release, and post-production monitoring.
  • Establish and manage SLOs/SLIs and error budgets; ensure they are integrated into roadmaps and delivery plans.
  • Lead DoD checklists and ensure automated tests, runbooks, and remediation steps are in place and validated.
  • Oversee operational readiness reviews and triage risks; perform root-cause analyses and drive auto-remediation.
  • Maintain logging, alerting, and monitoring platforms; govern CI/CD controls for security and reliability.
  • Lead incident response and post-incident reviews; drive resilience improvements and KPI monitoring.
  • Partner with engineering on cloud best practices and capacity planning for large-scale systems.

Skills

SRE leadership
Stakeholder management
Incident response
DevOps culture
Observability
Root-cause analysis
Cloud computing
Automation
Performance monitoring

Tools

Splunk
AppDynamics
Terraform
CloudFormation

Job description

Job Description

We are seeking a Delivery SRE leader who will ensure security applications are delivered with strong SDLC discipline and measurable reliability. This role partners closely with Product Owners and engineering leadership to challenge assumptions, sharpen the Definition of Done, and bake SRE requirements into design and build phases. The leader will govern operational readiness, quality gates, and resilience practices so that every release meets agreed SLOs and is production ready.

Key Responsibilities
  • Define and enforce quality gates across requirements, design, secure coding, testing, release, and post-production monitoring, translate business objectives into clear, testable requirements that include reliability, availability, performance, security, and observability.
  • Establish and manage SLOs/SLIs and error budgets; ensure they are integrated into product roadmaps and delivery plans, challenge Product Owners and teams to meet a rigorous, objective Definition of Done before release.
  • Sample DoD checklist: SLOs defined and monitored; alerts tuned; runbooks and escalation paths in place; automated tests (unit, integration, security) passing; performance and capacity validated; resilience and failover tested; rollback verified; vulnerability findings remediated; compliance controls and audit artifacts complete; documentation and support readiness confirmed.
  • Lead operational readiness reviews and triage risks; ensure timely remediation and prevention of recurrence through root-cause analysis and auto-remediation.
  • Maintain logging, alerting, and monitoring platforms; ensure dashboards provide health and performance visibility. Govern CI/CD pipeline controls for security, reliability, and change management; promote automation to eliminate toil.
  • Lead and participate in critical incident response (including outside business hours when needed); drive post-incident reviews and resilience improvements. Monitor delivery health and operational KPIs; lead continuous improvement across teams and products.
  • Oversee capacity planning and resilience management for large-scale, distributed systems, Partner with engineering on public cloud best practices (AWS or equivalent) for compute, storage, networking, messaging, automation (CloudFormation, Terraform), and data services.
  • Build a culture of collaboration, reliability, and continuous improvement; coach teams to adopt DevOps and SRE principles. Partner with regional engineering leaders to drive operational best practices and consistent execution. Provide concise, outcome-focused updates to management and stakeholders; influence decisions across Product, Engineering, SRE, and Security.
Required Qualifications, Capabilities, And Skills
  • Formal training or certification with 5+ years supporting critical security-focused applications in large-scale environments and managing and mentoring teams.
  • Experience with monitoring/logging tools (e.g., Splunk, AppDynamics) and dashboard technologies; Splunk Administrator certification desired.
  • Strong grasp of SDLC, secure development, DevOps/CI/CD tooling; capable of implementing top-tier continuous improvement with root-cause analysis and auto-remediation.
  • Effective under pressure; accountable, with excellent stakeholder management and communication skills.
  • This position may require HSA system access. Enhanced screening (criminal and credit background checks, and/or other screening) is required prior to employment and annually thereafter.
  • Global team collaboration with flexibility to engage during critical incidents outside standard business hours.
  • Experience implementing and managing SLOs/SLIs, error budgets, and operational readiness reviews for distributed systems, including leading post-incident analysis and resilience improvements.
  • Deep expertise in public cloud platforms (AWS or equivalent), infrastructure automation tools (CloudFormation, Terraform), and capacity planning for large-scale environments, with a track record of driving DevOps and SRE adoption across teams.
About Us

JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the worlds most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management.

We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.

We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants and employees religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.

JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans.

About The Team

Our professionals in our Corporate Functions cover a diverse range of areas from finance and risk to human resources and marketing. Our corporate teams are an essential part of our company, ensuring that we’re setting our businesses, clients, customers and employees up for success.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead SRE
Lead SRE

Next Frontier Capital • Plano (TX)

On-site
USD 180,000 - 240,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • Columbus (OH)

On-site
USD 150,000 - 190,000
SRE Software Engineer III
SRE Software Engineer III

JPMorganChase • Jersey City (NJ)

On-site
USD 140,000 - 180,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • New York (NY)

On-site
USD 120,000 - 160,000
Comprehensive health care coverage
Retirement savings plan
Tuition reimbursement
+2
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Fairygodboss • Jersey City (NJ)

On-site
USD 150,000 - 210,000
Site Reliability Engineer III
Site Reliability Engineer III

JPMorganChase • Plano (TX)

On-site
USD 130,000 - 180,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • Jersey City (NJ)

On-site
USD 170,000 - 230,000
AWS Certifications
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • Jersey City (NJ)

On-site
USD 170,000 - 250,000
Site Reliability Engineer III - AWS, Java and Kubernetes
Site Reliability Engineer III - AWS, Java and Kubernetes

Fairygodboss • Chicago (IL)

On-site
USD 120,000 - 150,000
Site Reliability Engineer III - AWS, Java and Kubernetes
Site Reliability Engineer III - AWS, Java and Kubernetes

JPMorganChase • Chicago (IL)

On-site
USD 130,000 - 190,000