Vice President - Lead Site Reliability Engineer

JPMorgan Chase

Plano (TX)

On-site

USD 150,000 - 230,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

JPMorganChase in Plano, TX is seeking a Lead Site Reliability Engineer to guide a team within the Infrastructure Platforms (Web Hosting) group, shaping reliability strategy and platform resilience.

You will lead incident response, apply AI-enabled tooling, and mentor peers while partnering with engineering and business stakeholders to improve service levels, scalability, and security.

Qualifications

  • Formal training or certification in site reliability engineering.
  • Proven ability to implement SRE best practices and toil reduction.

Responsibilities

  • Lead resiliency design reviews and mentor engineers.
  • Improve service levels with data-driven analysis.
  • Act as primary incident lead for critical apps.
  • Apply enterprise AI to incident triage and post-incident analysis.
  • Drive adoption of AI-assisted reliability workflows.
  • Collaborate with stakeholders to set SLOs and error budgets.
  • Ensure security, compliance, and auditability across platforms.

Skills

Python
Java/Spring Boot
Ansible
Observability
CI/CD
Networking basics
Mentoring

Education

SRE certification (e.g., SRE/DevOps)

Tools

Grafana
Dynatrace
Prometheus
CloudWatch
Splunk

Job description

Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.

As a Lead Site Reliability Engineer at JPMorganChase within the Infrastructure Platforms(Web Hosting) team, you hold a leadership role on your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business challenges they face. You will lead resiliency design reviews, break complex problems into digestible work for other engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to your peers.

In this role, you will partner with stakeholders across engineering and the business to drive reliability, stability, and continuous improvement across web hosting platforms. Your expertise will directly influence service levels, incident response, and the adoption of modern reliability practices - including the responsible use of enterprise-authorized AI capabilities - ensuring the firm's platforms remain resilient, scalable, and secure.

Job responsibilities
  • Consistently model and champion site reliability culture and practices, documenting and sharing knowledge across your organization through internal forums and communities of practice

  • Lead initiatives to improve the reliability and stability of web hosting platforms using data-driven analytics to improve service levels, proactively identifying and resolving technology-related bottlenecks within your areas of expertise

  • Drive collaboration with your team to identify comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgets with customers

  • Serve as the primary point of contact during major incidents for your applications, applying deep technical expertise to identify and resolve issues quickly to minimize business impact

  • Use enterprise-authorized AI capabilities to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data in accordance with sensitivity and security requirements

  • Lead reuse-first adoption of AI-assisted reliability workflows across software development lifecycle and toolchain practices - including continuous integration/continuous delivery quality checks, test and validation automation, and operational readiness - ensuring traceability, auditability, resiliency, and security controls

  • Provide ongoing guidance, tools, and solutions to support the firm's growth while working toward deep expertise on the applications and platforms within your scope, including their interdependencies and limitations

  • Offer mentorship and technical advice to other engineers, helping to grow site reliability knowledge and capability across the team

Required qualifications, capabilities, and skills
  • Formal training or certification on site reliability engineering concepts and 5+ years applied experience

  • Demonstrated proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices, with the ability to implement these within an application or platform

  • Advanced knowledge of site reliability culture and principles with demonstrated ability to apply them within an application or platform environment

  • Proficient knowledge and experience in observability, including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, CloudWatch, or Splunk

  • Demonstrated experience using enterprise-authorized AI capabilities to improve site reliability engineering workflows - such as incident investigation support and knowledge capture - with strong validation habits and awareness of data sensitivity

  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations

  • Fluency in at least one programming language such as Python, Java/Spring Boot, or Ansible

  • Proficient with continuous integration and continuous delivery practices and tooling, as well as container and container orchestration technologies

  • Experience troubleshooting common networking technologies and issues

  • Strong communication skills with the ability to mentor and educate others on site reliability principles and practices

Preferred qualifications, capabilities, and skills
  • Experience working with cloud platforms such as Amazon Web Services or Microsoft Azure, including an understanding of resiliency, scalability, and observability in cloud-native environments

  • Experience as a site reliability engineer supporting complex, mission-critical applications involving multiple components across varying technical generations

  • Exposure to artificial intelligence and machine learning concepts and their application in operational or reliability contexts

  • Demonstrated drive to self-educate, evaluate emerging technologies, and recommend suitable solutions that advance platform reliability

JPMorganChase, one of the oldest financial institutions, offers innovative financial solutions to millions of consumers, small businesses and many of the world's most prominent corporate, institutional and government clients under the J.P. Morgan and Chase brands. Our history spans over 200 years and today we are a leader in investment banking, consumer and small business banking, commercial banking, financial transaction processing and asset management.

We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.

We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants' and employees' religious practices and beliefs, as well as mental health or physical disability needs. Visit FAQs for more information about requesting an accommodation.

JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

Fairygodboss • Plano (TX)

On-site
USD 150,000 - 190,000
Comprehensive health care coverage
On-site health and wellness centers
Retirement savings plan
+4
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorganChase • Plano (TX)

On-site
USD 120,000 - 180,000
Health care coverage
On-site health & wellness centers
Retirement savings plan
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • Jersey City (NJ)

On-site
USD 170,000 - 230,000
AWS Certifications
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • Jersey City (NJ)

On-site
USD 170,000 - 250,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Fairygodboss • Plano (TX)

On-site
USD 120,000 - 150,000
Comprehensive health care coverage
Retirement savings plan
Tuition reimbursement
+1
Lead Site Reliability Engineer - Operations Excellence for AI Platforms
Lead Site Reliability Engineer - Operations Excellence for AI Platforms

JPMorganChase • Jersey City (NJ)

On-site
USD 140,000 - 200,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Worky • Plano (TX)

Hybrid
USD 150,000 - 210,000
comprehensive health care coverage
on-site health and wellness centers
retirement savings plan
+2
Senior Lead Site Reliability Engineer - AI/ML and Data Platforms
Senior Lead Site Reliability Engineer - AI/ML and Data Platforms

Next Frontier Capital • Jersey City (NJ)

On-site
USD 180,000 - 240,000
Director of Site Reliability Engineering
Director of Site Reliability Engineering

JPMorganChase • Jersey City (NJ)

On-site
USD 180,000 - 260,000
Competitive base salary
Discretionary incentive compensation
Health care coverage
+3
Site Reliability Engineer III
Site Reliability Engineer III

JPMorgan Chase • Jersey City (NJ)

On-site
USD 140,000 - 170,000
Comprehensive health care coverage
Retirement savings plan
Tuition reimbursement
+2