SRE Lead

Hdfc Securities

Mumbai

On-site

INR 3,500,000 - 5,500,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Hdfc Securities is seeking an SRE Lead to mentor a team of application support engineers, ensuring optimal system performance and uptime. You will oversee monitoring, troubleshooting, and maintenance of critical backend applications and infrastructure, and drive proactive strategies to minimize downtime and improve performance.

The role requires 8–10 years in Application Support or DevOps, with strong expertise in Linux scripting, Kubernetes, AWS, Terraform, Redis, and Oracle PL/SQL.

Qualifications

  • 8–10 years of hands-on experience in Application Support or DevOps roles.
  • Linux Shell Scripting: troubleshooting in Linux environments and automating tasks.
  • Kubernetes: managing and scaling Kubernetes clusters and deployments.
  • AWS Cloud: extensive knowledge of AWS services and CloudWatch logs.
  • Terraform: infrastructure as code for provisioning cloud resources.
  • Redis: database management, optimization, troubleshooting.
  • Oracle PL/SQL: querying, performance tuning, and data management.
  • CI/CD Tools: GitLab Runner, Argo CD, and pipelines.
  • API Gateway: configuring and managing secure API access.
  • ServiceNow/ITIL: incident, problem, and service request management.
  • Dynatrace/ELK/Grafana: APM, logging, dashboards, and alerts.
  • Incident management: best practices in high-availability environments.
  • Leading and mentoring support teams; strong problem-solving under pressure.
  • Educational requirement: B.E./B.Tech or M.E./M.Tech in CS/IT.

Responsibilities

  • Lead and mentor a team of application support engineers.
  • Oversee monitoring, troubleshooting, and maintenance of critical backend apps and infra.
  • Develop proactive support strategies to minimize downtime.
  • Coordinate with development teams to resolve complex issues and improve stability.
  • Oversee incident management and escalation procedures.
  • Prepare growth plans and perform regular performance evaluations.
  • Create runbooks, SOPs, and knowledge base articles.
  • Plan and execute extensive load testing to ensure reliability.

Skills

Leadership
Linux Shell Scripting
Kubernetes
AWS Cloud
Terraform
Redis
Oracle PL/SQL
CI/CD Tools
API Gateway
ITIL
Dynatrace
Grafana
Elasticsearch
Incident Management
Problem Solving

Education

B.E./B.Tech or M.E./M.Tech in CS/IT

Tools

GitLab Runner
Argo CD
Terraform (IaC)

Job description

As an SRE Lead, you will:
  • Lead, mentor, and guide a team of application support engineers, providing technical expertise and fostering a collaborative environment to ensure optimal system performance and uptime
  • Oversee the monitoring, troubleshooting, and maintenance of critical backend applications and infrastructure
  • Develop and implement proactive support strategies to minimize system downtime and improve overall application performance
  • Interact with stakeholders to understand operational requirements, analyze impact of changes, and execute improvements
  • Plan & execute extensive load testing for services to ensure system reliability under various conditions
  • Create and maintain runbooks, standard operating procedures (SOPs), and knowledge base articles for the support team
  • Analyze existing systems to identify improvements & performance / maintenance issues
  • Implement and manage monitoring and alerting systems to ensure early detection of potential issues
  • Coordinate with development teams to resolve complex application issues and contribute to the improvement of application stability
  • Oversee incident management processes, ensuring timely resolution of critical issues and proper escalation procedures
  • Prepare growth plans for the team and help them achieve their professional development goals
  • Provide regular feedback and performance evaluations for the team
We are looking for someone with:

8-10 years of hands-on experience in Application Support or DevOps roles

Linux Shell Scripting: Experience in troubleshooting issues on Linux environment. Proficient in writing and debugging shell scripts for automation and system management

Kubernetes: Deep experience in managing and scaling Kubernetes clusters and deployments. Experience in checking pod logs, number of nodes running, checking configurations etc.

AWS Cloud: Extensive knowledge of AWS services and architecture, including troubleshooting logs with CloudWatch

Terraform: Skilled in using Terraform for infrastructure as code, managing and provisioning cloud resources

Redis: Experience with Redis database management, optimization, and troubleshooting

Oracle PL/SQL: Strong expertise in Oracle PL/SQL for database querying, performance tuning, and data management in support contexts

CI/CD Tools: Hands-on experience with GitLab Runner, Argo CD, and other CI/CD tools for managing deployment processes and troubleshooting pipeline issues

API Gateway: Familiarity with configuring and managing API Gateways for secure API access and traffic management

ServiceNow/Similar tool with ITIL Framework: Practical experience with ServiceNow for incident management, problem management, and service request handling

Dynatrace/Elastic/Any other APM tool: Experience with Dynatrace for application performance monitoring, root cause analysis, and troubleshooting

Grafana: Proficiency in using Grafana for creating dashboards, visualizing performance metrics, and setting up alerts

Elasticsearch/Kibana: Knowledge of Elasticsearch for log analysis, searching, and creating visualizations in Kibana for operational insight

Deep understanding of incident management processes and best practices in a highly available environment

Proven experience in leading and managing support teams, providing technical guidance and mentorship

Strong problem-solving skills and ability to work under pressure in a fast-paced environment

Bachelor of Engineering (B.E) or Master's degree in Computer Science / Information Technology or equivalent experience

Following aspects would be a plus:
  • Prior experience in FinTech or other highly regulated industries
  • Experience with additional monitoring and observability tools
  • ITIL certification or strong knowledge of ITIL framework
  • Experience in performance tuning and optimization of large-scale applications
  • Familiarity with containerization technologies beyond Kubernetes (e.g., Docker)
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Lead
SRE Lead

Acldigital • Ahmedabad District

On-site
INR 1,500,000 - 2,000,000
Senior SRE Engineer
Senior SRE Engineer

Epam Systems • Bengaluru

On-site
INR 2,500,000 - 4,200,000
AWS SRE Professional
AWS SRE Professional

Infosys • Bengaluru

On-site
INR 900,000 - 1,300,000
Lead Support Analyst - Shared Services and Production Management , Information Technology
Lead Support Analyst - Shared Services and Production Management , Information Technology

CLSA • Pune District

On-site
INR 1,500,000 - 2,800,000
SRE Observability Engineer
SRE Observability Engineer

Awign • Hyderabad

On-site
INR 4,200,000 - 6,500,000
Senior SRE Technical Specialist
Senior SRE Technical Specialist

United States Digital Space LLC • Karnataka

On-site
INR 1,500,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

Innodata Inc. • India

On-site
INR 2,400,000 - 4,000,000
Site Reliability Engineer
Site Reliability Engineer

Epam Systems • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

Tekskills • Pune District

On-site
INR 1,200,000 - 1,800,000
Restaurant d'entreprise
Indemnités de stage/alternance
Site Reliability Engineer
Site Reliability Engineer

Tekskills • Chennai District

On-site
INR 1,800,000 - 3,000,000