Site Reliability Engineer II - Java/Python, Kubernetes, AWS, Terraform

JPMorgan Chase & Co.

Bengaluru

On-site

INR 1,800,000 - 3,000,000

Full time

7 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

JPMorgan Chase & Co. in Bengaluru is seeking an experienced Site Reliability Engineer II to improve system reliability and speed up incident handling.

You will apply software engineering best practices, collaborate across teams, and leverage enterprise AI capabilities to reduce toil and enhance observability. The role emphasizes building, maintaining, and improving SRE workflows, SLOs, and monitoring with tools like Grafana, Dynatrace, Prometheus, and Splunk, while using CI/CD pipelines and

Qualifications

  • Formal training or certification on site reliability engineering concepts and 2+ years applied experience.
  • Proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices.
  • Fluency in at least one programming language such as Java or Python.
  • Deep knowledge of software applications and technical processes with emerging depth in one or more technical disciplines.
  • Proficiency in observability and telemetry collection using Grafana, Dynatrace, Prometheus, Datadog, Splunk; CI/CD tooling like Jenkins, GitLab, Terraform.

Responsibilities

  • Executes small to medium projects independently with initial direction and graduates to designing and delivering projects independently.
  • Writes high quality, maintainable, and robust code following software engineering best practices.
  • Uses enterprise-authrozied AI capabilities to speed up incident triage, troubleshooting, and post-incident analysis.
  • Triages incidents and resolves problems at their root.
  • Eliminates toil through systems engineering or code updates.
  • Implements observability patterns and improves SLOs and alerting.
  • Applies AI capabilities to identify recurring toil and reliability risks from operational signals.
  • Experience with monitoring/logging tools and dashboards (Splunk, AppDynamics, Grafana).
  • Experience with container orchestration (Kubernetes, ECS, Docker).
  • Troubleshooting common networking technologies and issues.

Skills

Java/Python
Observability/Monitoring
CI/CD/DevOps
Kubernetes/Docker
SRE Fundamentals
Security & SDLC
AI in SRE
Troubleshooting Networks
Incident Response
Automation/Toil Reduction

Tools

Grafana
Dynatrace
Prometheus
Datadog
Splunk
Jenkins
GitLab CI
Terraform
Kubernetes
Docker
AppDynamics

Job description

Play a key role in ensuring system reliability at one of the world’s most iconic and largest financial institutions.

As a Site Reliability Engineer II at JPMorgan Chase within the Commercial Investment Bank team, you will use technology to solve business problems and leverage software engineering best practices as we strive towards excellence. This role often works independently to execute small to medium projects, but you’ll also have the opportunity to collaborate with cross functional teams to continually improve your level of knowledge about JPMorgan Chase’s business and relevant technologies.

Job responsibilities
  • Executes small to medium projects independently with initial direction and graduates to designing and delivering projects independently
  • Leverages technology to solve business problems by writing high quality, maintainable, and robust code following best practices in software engineering
  • Uses enterprise-authorized AI capabilities within the work environment to speed up incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Participates in triaging, examining, diagnosing, and resolving incidents and works with others to solve problems at their root
  • Recognizes toil within the role and proactively works towards eliminating it through systems engineering or updating application code
  • Understands observability patterns and strives to implement and improve service level indicators, objectives monitoring, and alerting solutions for optimal transparency and analysis
  • Applies enterprise-authorized AI capabilities within the work environment to identify recurring toil and reliability risks from operational signals, prioritizing reuse-first improvements and measurable SLO outcomes.
Required qualifications, capabilities, and skills
  • Formal training or certification on site reliability engineering concepts and 2+ years applied experience
  • Proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices with the ability to implement these practices within an application or platform
  • Fluency in at least one programming language such as (e.g., Java/Python, Java Spring Boot, etc.)
  • Deep knowledge of software applications and technical processes with emerging depth in one or more technical disciplines
  • Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, etc. Proficiency in continuous integration and continuous delivery tools (e.g., Jenkins, GitLab, Terraform, etc.)
  • Grasp of SDLC, secure development, DevOps/CI/CD tooling; capable of implementing top-tier continuous improvement with root-cause analysis and auto-remediation.
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
  • Use evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
  • Experience with monitoring/logging tools (e.g., Splunk, AppDynamics) and dashboard technologies.
  • Experience with container and container orchestration (e.g., ECS, Kubernetes, Docker, etc.)
  • Experience with troubleshooting common networking technologies and issues

Preferred qualifications, capabilities, and skills
  • Tools / Agentic AI to solve common opportunities in SRE Domain
  • Certification towards CI/CD and AI skills. Splunk Administrator certification desired.
  • Drive to self-educate and evaluate new technology

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II - Java/Python, Kubernetes, AWS, Terraform
Site Reliability Engineer II - Java/Python, Kubernetes, AWS, Terraform

JPMorganChase • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer II
Site Reliability Engineer II

Next Frontier Capital • Bengaluru

On-site
INR 4,500,000 - 6,500,000
Site Reliability Engineer
Site Reliability Engineer

JP Morgan Services India Pvt Ltd • Bengaluru

On-site
INR 1,500,000 - 2,300,000
Lead Site Reliability Engineer AWS
Lead Site Reliability Engineer AWS

JPMorganChase • Bengaluru

On-site
INR 3,500,000 - 5,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

JP Morgan Services India Pvt Ltd • Bengaluru

On-site
INR 3,000,000 - 4,200,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Next Frontier Capital • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Fairygodboss • Mumbai

On-site
INR 3,500,000 - 6,000,000
Site Reliability Engineer II
Site Reliability Engineer II

JPMorganChase • Bengaluru

On-site
INR 2,000,000 - 3,200,000
Site Reliability Engineer II
Site Reliability Engineer II

JPMorgan Chase & Co. • Bengaluru

On-site
INR 1,800,000 - 3,000,000
Software Engineer II - Python & AWS
Software Engineer II - Python & AWS

JPMorganChase • Hyderabad

On-site
INR 1,400,000 - 2,100,000