Tech & Digital-Lead Site Reliability Engineer

Hdfc Bank

Bengaluru

On-site

INR 3,500,000 - 5,500,000

Full time

6 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Hdfc Bank seeks a Lead Site Reliability Engineer to analyze, troubleshoot, and design reliable services on GCP with a focus on scalability, resilience, security, and performance.

Lead and mentor a team of SRE engineers, driving reliability initiatives and improving observability and incident response across a cloud and on-prem infrastructure.

Qualifications

  • B Tech in Computer Science or related discipline preferred.
  • 11–13 years of total experience as an SRE/DevOps professional.
  • Hands-on with cloud platforms (GCP preferred) and automation.
  • Strong automation, observability, and incident management skills.

Responsibilities

  • Execute reliability initiatives for the team and organization.
  • Mеntor and lead a team of SRE engineers.
  • Help build an SRE culture by sharing practices and code across teams.
  • Solid understanding of observability tools and reliability metrics.
  • Define KPI in terms of RPO/RTO/SLI/SLO/Error Budget.
  • Apply automation to manual tasks and systems.
  • Troubleshoot cross-platform issues and manage live incidents.
  • Monitor performance and improve overall stability and reliability.
  • Conduct system analysis and develop improvements for performance.
  • Design, write, and ship software to increase observability and efficiency.
  • Maintain deployment orchestration of servers, containers, and databases.
  • Develop run books and SOPs for recurring production issues.

Skills

Monitoring infra
Observability
Scripting
Containerization
Orchestration
Programming basics
Observability tools

Education

BTech in Computer Science

Tools

Docker
Kubernetes
Terraform
CloudFormation
Ansible
Kafka

Job description

Job Title

Lead Site Reliability Engineer

Job Details
  • Business Unit: Tech & Digital
  • Team: Ent Factory-Channels, Mobility, Payments
  • Reports to: SRE Manager
  • Location: Mumbai, Chennai, Gurgaon & Bangalore
  • Role Type: Non-Supervisory
  • No of direct reportees: 0
  • Travel Required: No
  • Job Band Range: D1/D2
  • JD Created date: 28th Jan 2023
Job Purpose

Analyzing, troubleshooting, and designing vital services, platforms, and infrastructure on GCP with a focus on reliability, scalability, resilience, security, and performance.

Lead and Mentor a team of SRE engineers.

Job Responsibilities
  • Execute reliability initiatives for the team and organization.
  • Mentor and lead a team of SRE engineers.
  • Help build a Site Reliability Engineering culture by sharing best practices, approaches, documentation, and code with other engineering teams.
  • Solid understanding of observability tools and ability to express reliability metrics via observability.
  • Define KPI in the form of RPO/RTO/SLI/SLO/Error Budget.
  • Apply automation and software to any manually performed tasks or system parts.
  • Troubleshoot complicated, cross-platform issues handling OS, Networking, Database in a cloud-based SaaS environment and manage live production incidents.
  • Monitor application performance, take steps to improve overall application performance and stability, and follow through with implementation.
  • Conduct system analysis, configuration management, and develop improvements for system software performance, availability, and reliability.
  • Design, write, ship, and motivate the creation of software and systems to increase observability, product reliability, and organizational efficiency.
  • Maintain and monitor deployment, orchestration of servers, docker containers, databases, and general backend infrastructure.
  • Develop Run Books/Standard Operating Procedure for recurring Production issues and work on permanent solutions.
  • Perform Incident Analysis regularly to prevent and find long-term solutions for Incidents.
Educational Qualifications

B Tech in Computer Science or related discipline preferred.

Key Skills
  • Experience in monitoring and analyzing infrastructure performance using standard performance monitoring tools.
  • Demonstrable experience in Containerization (Docker) and orchestration (Kubernetes).
  • Experience with Infrastructure As Code (Terraform, Cloud Formation, Ansible).
  • Knowledge and proven hands-on experience in large-scale databases and distributed technologies, such as Kafka and Confluent Platform Kafka.
  • Basic programming and scripting skills.
  • Solid understanding of at least 2 observability technologies.
Experience Required

Total Years of experience: 11-13

Major Stakeholders

Internal: Product Manager from Digital Factory, Business Analyst from BTG team, Incident Management team, Development Team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hdfc Bank • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer
Site Reliability Engineer

Mumba Technologies, Inc. • Gurugram District

Hybrid
INR 2,500,000 - 4,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

VMC Soft Technologies, Inc • Hyderabad

Hybrid
INR 1,500,000 - 2,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Tata Consultancy Services • Bengaluru

On-site
INR 2,800,000 - 4,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hilabs • Bengaluru

On-site
INR 2,200,000 - 3,500,000
Site Reliability Engineer Lead Position
Site Reliability Engineer Lead Position

Cloudxtreme • Hyderabad

On-site
INR 2,600,000 - 5,000,000
Site Reliability Engineer
Site Reliability Engineer

Zorba AI • Chennai District

On-site
INR 4,000,000 - 7,500,000
Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Maharashtra

On-site
INR 700,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

Zorba AI • Hyderabad

On-site
INR 4,000,000 - 7,500,000
Site Reliability Engineer (SRE) – GCP Platform
Site Reliability Engineer (SRE) – GCP Platform

ITC Infotech • Bengaluru

On-site
INR 900,000 - 1,300,000