SRE: Reliability, Incidents & Cloud Observability

JPMorganChase

Plano (TX)

On-site

USD 120,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

JPMorganChase is seeking a Senior Site Reliability Engineer to define reliability targets and architect robust monitoring across complex environments. You will lead incident management, automate recovery workflows, and ensure high availability across multi-region cloud deployments.

The role requires deep expertise in Kubernetes, AWS, databases, and security, with a focus on reducing MTTR and preventing outages through proactive observability and robust runbooks.

Qualifications

  • Bachelor's degree in Information Systems Engineering, Computer Engineering, or related field plus 5 years of experience in the job offered or as Site Reliability Engineer or related roles.
  • Five years of experience in Sev1/Sev2 on-call, MTTR reduction, observability, automation, CI/CD, Terraform, Kubernetes, and AWS production operations.
  • Three years of experience with SLI/SLO, change governance, performance testing, DR/reliability, and scalable systems.

Responsibilities

  • Define and enforce reliability targets for critical environments with measurable metrics.
  • Architect, implement, and refine monitoring and alerting across complex stacks.
  • Lead incident management, root cause analysis, and long-term improvements.
  • Automate failover and recovery workflows for services across cloud regions.
  • Maintain runbooks, dashboards, and pre/post deployment validation to prevent drift.

Skills

Sev1/Sev2 on-call
Observability
Automation Python Bash
CI/CD blue/green canary
Terraform
Kubernetes
AWS production ops
PostgreSQL
MySQL
Oracle backups
Linux networking
Security: least privilege
SLI/SLO
Gate changes
Performance testing
JMeter BlazeMeter
Reliability engineering

Education

Bachelor's degree in Information Systems Engineering/Computer Engineering

Tools

Terraform

Job description

JPMorganChase is seeking a Senior Site Reliability Engineer to define reliability targets and architect robust monitoring across complex environments. You will lead incident management, automate recovery workflows, and ensure high availability across multi-region cloud deployments.

The role requires deep expertise in Kubernetes, AWS, databases, and security, with a focus on reducing MTTR and preventing outages through proactive observability and robust runbooks.

Get your free, confidential resume review.
or drag and drop your file here.