Lead Site Reliability Engineer

Fairygodboss

Mumbai

On-site

INR 3,500,000 - 6,000,000

Full time

12 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

JPMorgan Chase in Mumbai seeks an experienced Lead Site Reliability Engineer to drive reliability across AI/ML infrastructure. You will lead the SRE team, set standards for monitoring, incident response, and proactive improvements, collaborating with stakeholders to define SLIs/SLOs and improve service levels.

The role requires deep cloud, automation, and observability expertise with hands-on leadership, mentoring, and a focus on security and resiliency within enterprise systems.

Qualifications

  • 8+ years in site reliability or infrastructure roles.
  • Experience with AWS, Terraform, and monitoring tools like Prometheus, Grafana, CloudWatch.
  • Strong problem-solving, communication, and collaboration skills including CI/CD and risk management.
  • Bachelor's or Master's degree in Computer Science, Engineering, or related field (or equivalent).
  • Experience leading a team in enterprise AI reliability workflows, including validation and data sensitivity awareness.
  • Fluency in Python and knowledge of software processes with emerging depth in technical domains.
  • Proficiency in observability and SRE practices using tools like Grafana, Dynatrace, Prometheus, Datadog, Splunk.
  • Proficiency in CI/CD tools (Jenkins, GitLab, Terraform) and container orchestration (ECS, Kubernetes, Docker).
  • Experience with networking troubleshooting and continuous learning.

Responsibilities

  • Design, implement, and optimize SRE practices for AI/ML infrastructure focusing on reliability and scalability.
  • Develop and maintain automated monitoring, alerting, and incident response systems.
  • Manage day-to-day support issues, conduct root cause analysis, and drive continuous improvement to reduce repeats.
  • Champion SRE culture and provide technical influence across the team.
  • Lead initiatives to improve reliability using data-driven analytics and improve service levels.
  • Collaborate to identify SLIs/SLOs and establish error budgets with stakeholders.
  • Drive adoption of enterprise AI capabilities with proper governance and data handling.
  • Demonstrate technical expertise and solve bottlenecks in key domains.
  • Act as main point of contact during major incidents to minimize financial impact.
  • Establish reuse-first practices for AI workflows and ensure traceability of changes.
  • Document and share knowledge via internal forums and communities of practice.

Skills

SRE leadership
Python
problem solving
communication
AI reliability

Education

Bachelor's or Master's in Computer Science / Engineering

Tools

AWS
Terraform
CloudFormation
Prometheus
Grafana
CloudWatch
Jenkins
GitLab CI
Kubernetes
Docker
Dynatrace
Datadog
Splunk
ECS

Job description

Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability.

As a Lead Site Reliability Engineer at JPMorgan Chase within Corporate Technology, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them.

Job responsibilities
  • Design, implement, and optimize SRE best practices for AI/ML infrastructure, focusing on reliability, scalability, security, and operational efficiency.
  • Develop and maintain automated monitoring, alerting, and incident response systems.
  • Manage day-to-day support issues, conduct root cause analysis, and drive continuous improvement to reduce repeat errors and enhance system stability.
  • Demonstrates and champions site reliability culture and practices and exerts technical influence throughout your team
  • Leads initiatives to improve the reliability and stability of your team's applications and platforms using data-driven analytics to improve service levels
  • Collaborates with team members to identify comprehensive service level indicators and stakeholders to establish reasonable service level objectives and error budgets with customers
  • Drives team adoption of enterprise-authorized AI capabilities within the work environment to improve reliability delivery speed and operational efficiency (e.g., drafting/runbook hygiene, incident learning capture, and backlog prioritization), with human-in-the-loop validation and appropriate handling of sensitive data.
  • Demonstrates a high level of technical expertise within one or more technical domains and proactively identifies and solves technology-related bottlenecks in your areas of expertise
  • Acts as the main point of contact during major incidents for your application and demonstrates the skills to identify and solve issues quickly to avoid financial losses
  • Establishes reuse-first practices for AI-assisted workflows across team delivery and operational routines, reinforcing resiliency and security expectations and ensuring traceability/auditability of changes.
  • Documents and shares knowledge within your organization via internal forums and communities of practice
Required qualifications, capabilities, and skills
  • 8+ years in site reliability or infrastructure engineering roles.
  • Deep expertise in AWS cloud services, infrastructure automation (Terraform, CloudFormation), and monitoring tools (Prometheus, Grafana, CloudWatch).
  • Strong problem-solving, communication, and collaboration skills. Experience with CI/CD pipelines, operational stability, and risk management.
  • Bachelor's or Master's degree in Computer Science, Engineering, or related field (or equivalent experience).
  • Experience leading a team in the safe use of enterprise-authorized AI capabilities within the work environment for reliability engineering workflows, including validation habits and awareness of data sensitivity. Ability to set team-level expectations for reviewing AI-assisted recommendations and escalating uncertain decisions while maintaining resiliency, security, and auditability outcomes. Deep proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices with the ability to implement these practices within an application or platform
  • Fluency in at least Python programming language and deep knowledge of software applications and technical processes with emerging depth in one or more technical disciplines
  • Proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, etc.
  • Proficiency in continuous integration and continuous delivery tools (e.g., Jenkins, GitLab, Terraform, etc.)
  • Experience with container and container orchestration (e.g., ECS, Kubernetes, Docker, etc.)
  • Experience with troubleshooting common networking technologies and issues and drive to self-educate and evaluate new technology
  • Strong communication skills with ability to mentor and educate others on site reliability principles and practices
Preferred qualifications, capabilities, and skills
  • Certified in SRE
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorganChase • Mumbai

On-site
INR 4,200,000 - 6,600,000
Lead Site Reliability Engineer AWS
Lead Site Reliability Engineer AWS

JPMorganChase • Bengaluru

On-site
INR 3,500,000 - 5,200,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Next Frontier Capital • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer II - Java/Python, Kubernetes, AWS, Terraform
Site Reliability Engineer II - Java/Python, Kubernetes, AWS, Terraform

JPMorgan Chase & Co. • Bengaluru

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer II - Java/Python, Kubernetes, AWS, Terraform
Site Reliability Engineer II - Java/Python, Kubernetes, AWS, Terraform

JPMorganChase • Bengaluru

On-site
INR 1,200,000 - 2,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

JP Morgan Services India Pvt Ltd • Bengaluru

On-site
INR 3,000,000 - 4,200,000
Site Reliability Engineer III
Site Reliability Engineer III

JPMorganChase • Mumbai

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer II
Site Reliability Engineer II

Next Frontier Capital • Bengaluru

On-site
INR 4,500,000 - 6,500,000
Site Reliability Engineer
Site Reliability Engineer

JP Morgan Services India Pvt Ltd • Bengaluru

On-site
INR 1,500,000 - 2,300,000
Site Reliability Engineer III
Site Reliability Engineer III

JPMorgan Chase & Co. • Hyderabad

On-site
INR 1,500,000 - 2,000,000