Site Reliability Engineer II - Python, Observability, AWS, Terraform

JPMorganChase

Mumbai

On-site

INR 1,800,000 - 2,400,000

Full time

13 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

JPMorgan Chase is seeking a Site Reliability Engineer II within the Commercial & Investment Bank Payments Technology team in Mumbai. The role focuses on reliability, observability, and automation, collaborating with cross‑functional teams to improve systems and reduce toil.

The candidate will apply cloud skills, AI tooling for incident triage, and strong coding capabilities to build robust, scalable services while maintaining security and auditability in a high‑stakes financial environment.

Qualifications

  • 2+ years of applied software engineering experience.
  • Proficient in site reliability culture and SLI/SLO/SLA concepts.
  • Experience with observability and telemetry tools (Grafana, Dynatrace, Prometheus, Datadog, Splunk).
  • Strong knowledge of software domains (Cloud, AI, Android, etc.) with hands-on system design, resiliency, testing and disaster recovery.
  • Strong AWS skills (EC2, S3, RDS, VPC, IAM) and networking knowledge.
  • Proficient in Python or Java/Spring Boot for automation and tooling to reduce toil.
  • Experience with CI/CD tooling and running production incident calls.
  • Experience with Jenkins, GitLab, or Terraform for CI/CD pipelines.
  • Cloud production support in AWS with debugging distributed systems.

Responsibilities

  • Executes small to medium projects independently with initial direction; grows to independent delivery.
  • Writes high quality, maintainable code following software engineering best practices.
  • Triages and resolves incidents; collaborates with others to solve root causes.
  • Identifies toil and eliminates it via system engineering or code improvements.
  • Implements observability patterns, SLIs/SLAs, monitoring, and alerting for visibility.
  • Uses enterprise AI capabilities to speed up triage, troubleshooting, and post-incident analysis while ensuring data sensitivity.
  • Assesses AI-assisted recommendations for risk and resiliency; applies necessary controls.

Skills

SLI/SLO/SLA
Error budgets
AWS
Python/Java
CI/CD tooling
Incident management
Logging/telemetry
Monitoring tools
Distributed systems
Observability

Tools

Grafana
Dynatrace
Prometheus
Datadog
Splunk
Geneos
Jenkins
GitLab
Terraform
Kubernetes

Job description

Job Description

Play a key role in ensuring system reliability at one of the world's most iconic and largest financial institutions.

Job Description

Play a key role in ensuring system reliability at one of the world's most iconic and largest financial institutions. As a Site Reliability Engineer II at JPMorgan Chase within the Commercial & Investment Bank Payments Technology team, you will use technology to solve business problems and leverage software engineering best practices as we strive towards excellence. This role often works independently to execute small to medium projects, but you'll also have the opportunity to collaborate with cross functional teams to continually improve your level of knowledge about JPMorgan Chase's business and relevant technologies.

Job Responsibilities
  • Executes small to medium projects independently with initial direction and eventually graduates to designing and delivering projects by yourself
  • Leverages technology to solve business problems by writing high quality, maintainable, and robust code following best practices in software engineering
  • Participates in triaging, examining, diagnosing, and resolving incidents and work with others to solve problems at their root
  • Recognizes the toil within your role and proactively works towards eliminating it through either systems engineering or updating application code
  • Understands observability patterns and strives to implement and improve service level indicators, objectives monitoring, and alerting solutions for optimal transparency and analysis
  • Uses enterprise-authorized AI capabilities within the work environment to speed up incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements
  • Applies enterprise-authorized AI capabilities within the work environment to identify recurring toil and reliability risks from operational signals, prioritizing reuse-first improvements and measurable SLO outcomes
Required Qualifications, Capabilities, And Skills
  • Formal training or certification on software engineering concepts and 2+ years applied experience
  • Proficient in site reliability culture and principles and familiarity with how to implement site reliability within an application or platform, including strong understanding of SLI/SLO/SLA and error budgets
  • Experience in observability such as white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Geneos, Prometheus, Datadog, Splunk, and others
  • Proficient knowledge of software applications and technical processes within a given technical discipline (e.g., Cloud, AI, Android, etc.), including hands‑on experience in system design, resiliency, testing, operational stability, and disaster recovery
  • Strong expertise in AWS services, including EC2, S3, RDS, VPC, IAM, and networking.
  • Proficient in at least one programming language such as Python, Java/Spring Boot for AI/ML modeling and automation to reduce operational toil by building tools for repeated tasks
  • Experience with continuous integration and continuous delivery tooling, along with experience running production incident calls and managing incident resolution in collaboration with cross-functional teams
  • Experience with continuous integration and continuous delivery tools like Jenkins, GitLab, or Terraform
  • Practical cloud experience in AWS supporting production services (visibility/troubleshooting, deployments, and operational hygiene) and strong debugging fundamentals across distributed systems (logs/metrics, latency analysis, dependency failures, data issues)
  • Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows (e.g., troubleshooting support and runbook drafting) with strong validation habits and awareness of data sensitivity. Ability to assess AI-assisted operational recommendations for correctness and risk, and apply appropriate controls to maintain resiliency, security, and auditability
Preferred Qualifications, Capabilities, And Skills
  • Experience with continuous integration and continuous delivery tooling
  • Experience applying AI-assisted tooling to reduce support toil
  • Familiarity with container and container orchestration and troubleshooting common networking technologies and issues
  • Ability to identify new technologies and relevant solutions to ensure design constraints are met by the software team
  • Ability to initiate and implement ideas to solve business problems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II
Site Reliability Engineer II

JPMorgan Chase & Co. • Bengaluru

On-site
INR 1,000,000 - 1,800,000
Site Reliability Engineer III
Site Reliability Engineer III

JPMorgan Chase & Co. • Hyderabad

On-site
INR 1,500,000 - 2,000,000
Lead Software Engineer - DevOps/SRE/AWS/EKS
Lead Software Engineer - DevOps/SRE/AWS/EKS

JPMorganChase • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Software Engineer III - Python, Java, Observability, OTel, Dynatrace
Software Engineer III - Python, Java, Observability, OTel, Dynatrace

JPMorganChase • Mumbai

On-site
INR 1,200,000 - 2,000,000
Lead Software Engineer - Java, Spring boot, Microservices , Real time application
Lead Software Engineer - Java, Spring boot, Microservices , Real time application

JPMorganChase • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Software Engineer Iii - Sre, Python/Java, Aws, Ai
Software Engineer Iii - Sre, Python/Java, Aws, Ai

JPMorganChase • Telangana

On-site
INR 1,800,000 - 2,800,000
Lead Software Engineer - (SRE principles , AI tools, AWS)
Lead Software Engineer - (SRE principles , AI tools, AWS)

JPMorgan Chase & Co. • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Lead Software Engineer - Cloud and AI-Assisted Development
Lead Software Engineer - Cloud and AI-Assisted Development

JP Morgan Services India Pvt Ltd • Bengaluru

On-site
INR 3,000,000 - 5,400,000
Software Engineer III
Software Engineer III

JPMorgan Chase Bank • Mumbai

On-site
INR 1,800,000 - 2,600,000
Lead Software Engineer
Lead Software Engineer

JPMorganChase • Mumbai

On-site
INR 2,500,000 - 3,500,000