Sr Lead Site Reliability Engineer - Cloud Data & Databricks

JPMorgan Chase Bank

Hyderabad

On-site

INR 600,000 - 1,200,000

Full time

12 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

JPMorgan Chase & Co. is seeking a Senior Lead Site Reliability Engineer in Hyderabad to lead a team focused on reliable, scalable data platforms.

You will manage AWS Databricks deployments, drive observability, and collaborate with cross-functional data teams to accelerate reliability decisions. You will apply AI-enabled reliability workflows across the SDLC, ensure security controls, and mentor engineers in best practices while delivering high-quality production code and robust incident

Qualifications

  • Formal training or certification related to software engineering concepts and 10+ years of applied experience.

Responsibilities

  • Lead a team of SREs to design, implement, and maintain a managed AWS Databricks platform.
  • Perform platform design, setup and configuration, workspace administration, and resource monitoring.
  • Create multi-AZ, multi-region, and multi-cloud resiliency strategies for business-critical products and services.
  • Lead evaluation sessions with external vendors, startups, and internal teams to assess architectural designs and technical credentials.
  • Drive continuous improvement in system observability, alerting, and capacity planning.
  • Collaborate with engineering and data teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence.
  • Execute creative software solutions, design, development, and technical troubleshooting.
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices, ensuring traceability/auditability, resiliency, and security controls.
  • Develop secure, high-quality production code, and review/debug code written by others to ensure reliability and correctness.
  • Apply SRE best practices to improve reliability, scalability, and performance; eliminate or automate recurring issues; and maintain incident response procedures including root cause analysis and postmortems.

Skills

SRE Principles
SLIs/SLOs
Incident management
CI/CD pipelines
Python
AI in ops
Terraform
Spark
AWS Databricks
Docker
Kubernetes
Big data
Troubleshooting
Communication

Tools

Databricks Platform
Terraform
Docker
Kubernetes

Job description

Join us to shape the future of data and analytics, leveraging your expertise to deliver impactful technology solutions. Experience career growth and make a difference in a collaborative environment.

As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the AIML Data Platforms and Chief Data and Analytics Team, you will develop and deliver advanced technology products focused on data and analytics. You will tackle complex cloud data platform challenges, especially around Data Lake Tools, and work in an agile environment, collaborating with cross-functional teams. You will help drive the firm s data and analytics journey, ensuring quality, integrity, and security of data, and leveraging AI/ML technologies to support commercial goals. You will contribute to a culture of innovation and operational excellence.

Job responsibilities
  • Lead a team of SREs to design, implement, and maintain a managed AWS Databricks platform, providing engineering and operational support to Application/Engineering teams
  • Perform platform design, set-up and configuration, workspace administration, and resource monitoring
  • Create multi-AZ, multi-region, and multi-cloud resiliency strategies for business-critical products and services
  • Lead evaluation sessions with external vendors, startups, and internal teams to assess architectural designs and technical credentials
  • Drive continuous improvement in system observability, alerting, and capacity planning
  • Collaborate with engineering and data teams to optimize infrastructure and deployment processes, focusing on automation and operational excellence
  • Execute creative software solutions, design, development, and technical troubleshooting
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate reliability design and operational decisioning (e.g., incident/post-incident analysis and requirements traceability), validating outputs and handling operational data according to sensitivity and security requirements.
  • Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., testing/validation automation and production readiness), ensuring traceability/auditability, resiliency, and security controls.
  • Develop secure, high-quality production code, and review/debug code written by others to ensure reliability and correctness.
  • Apply SRE best practices to improve reliability, scalability, and performance; eliminate or automate recurring issues; and maintain incident response procedures including root cause analysis and postmortems.
Required qualifications, capabilities and skills
  • Formal training or certification on software engineering concepts and 10+ years applied experience
  • Strong understanding of SRE principles, including SLIs, SLOs, error budgets, and incident management
  • Experience with monitoring tools, automation frameworks, and CI/CD pipelines
  • Proficient in Python application program development with use of automated unit testing
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve reliability engineering workflows with strong validation habits and awareness of data sensitivity.
  • Ability to set team practices for safe AI usage in operations (e.g., review/approval expectations and escalation paths) while maintaining resiliency, security, and auditability outcomes.
  • Experience with Terraform development and understanding of Terraform enterprise
  • Experience in delivering system design, application development, testing, and operational stability
  • Knowledge of Big Data distributed compute frameworks like Spark, Glue, MapReduce
  • Excellent troubleshooting, analytical, and communication skills
Preferred qualifications, capabilities and skills
  • Experience in Data pipelines using Spark
  • Exposure to AWS & Databricks Platform administration
  • Knowledge of containerization (Docker, Kubernetes) and orchestration
  • Familiarity with distributed systems and large-scale data processing

Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Lead Software Engineer Java FSD + AWS
Senior Lead Software Engineer Java FSD + AWS

JPMorgan Chase & Co. • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Senior Lead Software Engineer Java FSD AWS
Senior Lead Software Engineer Java FSD AWS

JPMorganChase • Bengaluru

On-site
INR 3,500,000 - 7,500,000
Senior Lead Software Engineer - Java Full Stack Developer with AWS
Senior Lead Software Engineer - Java Full Stack Developer with AWS

JP Morgan Services India Pvt Ltd • Bengaluru

On-site
INR 3,600,000 - 6,000,000
Senior Lead Software Engineer Java FSD + AWS
Senior Lead Software Engineer Java FSD + AWS

JPMorganChase • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Lead Software Engineer - Python, AWS, BigData
Lead Software Engineer - Python, AWS, BigData

Fairygodboss • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Senior Lead Software Engineer - Java & Databricks
Senior Lead Software Engineer - Java & Databricks

JPMorgan Chase & Co. • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Lead Software Engineer - Java, Spring boot, Microservices , Real time application
Lead Software Engineer - Java, Spring boot, Microservices , Real time application

Fairygodboss • Hyderabad

On-site
INR 1,000,000 - 1,500,000
Lead Software Engineer
Lead Software Engineer

JPMorgan Chase & Co. • Bengaluru

On-site
INR 2,600,000 - 3,800,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

JPMorgan Chase & Co. • Mumbai

On-site
INR 4,000,000 - 6,000,000
Sr Lead Software Engineer - Java & AI
Sr Lead Software Engineer - Java & AI

JPMorgan Chase Bank • Hyderabad

On-site
INR 4,000,000 - 7,000,000