Site Reliability Engineer Lead

TymblHub

India

On-site

INR 2,500,000 - 4,000,000

Full time

40 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Good co India is seeking a Site Reliability Engineer Lead to own the reliability of critical production systems. You will establish SRE practices, lead incident response, and drive automation across cloud-native environments.

The role requires strong Linux skills and hands-on experience with Docker, Kubernetes, Terraform, and Ansible. Collaboration with product, security, and infra teams is essential to improve availability and performance.

Qualifications

  • 5–10 years of experience in SRE/DevOps or related field.
  • Strong hands-on in production systems and large-scale distributed apps.
  • Proficient in Linux and scripting languages (Python, Go, Bash).

Responsibilities

  • Lead the SRE function and ensure availability, reliability, and performance of critical apps.
  • Define and implement SRE practices with SLIs, SLOs, SLAs and error budgets.
  • Design, operate and scale cloud-native systems across multi-cloud or hybrid environments.
  • Lead incident response, RCA, and post-incident reviews for production issues.
  • Drive automation of infrastructure and release processes; manage CI/CD pipelines.

Skills

SRE
DevOps
Cloud Engineering
Linux
Python

Tools

Docker
Kubernetes
Terraform
Ansible
Prometheus

Job description

# Site Reliability Engineer LeadGood co IndiaPosted on October 2, 2026---Experience5 - 10 yrsSalary (CTC)₹25L - ₹40LJob LocationIndiaVacancy1DesignationSite Reliability Engineer LeadJob TypeNot specified---## Job Description**Role & Responsibilities*** Lead the Site Reliability Engineering function and ensure the availability, reliability, scalability, and performance of critical applications and infrastructure.* Define and implement SRE practices around SLIs, SLOs, SLAs, error budgets, and service reliability.* Design and operate highly available, fault-tolerant, and scalable cloud-native systems.* Lead incident response, troubleshooting, root-cause analysis, and post-incident reviews for critical production issues.* Build and improve monitoring, observability, logging, alerting, and performance-management capabilities.* Drive automation of infrastructure, deployment, operational, and repetitive engineering processes.* Manage and optimize cloud infrastructure across AWS, Azure, GCP, or hybrid environments.* Implement Infrastructure as Code using tools such as Terraform and Ansible.* Manage containerized and orchestration platforms including Docker and Kubernetes.* Establish and improve CI/CD pipelines, deployment strategies, release automation, and rollback mechanisms.* Lead capacity planning, performance optimization, disaster recovery, backup, and business-continuity initiatives.* Identify and reduce operational risks, technical debt, toil, and recurring production issues.* Implement reliability and resilience engineering practices, including failure testing and chaos engineering where appropriate.* Partner with development, infrastructure, security, and product teams to improve application reliability throughout the software lifecycle.* Establish SRE standards, operational runbooks, documentation, and engineering best practices.* Mentor SRE/DevOps engineers and provide technical leadership across reliability initiatives.* Track infrastructure and cloud costs and identify opportunities for performance and cost optimization. **Preferred Candidate Profile*** 5 to 10 years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, or a related field.* Strong hands-on experience managing production systems and large-scale, distributed applications.* Strong proficiency in Linux and scripting/programming languages such as Python, Go, or Bash.* Extensive experience with AWS, Azure, GCP, or hybrid-cloud environments.* Strong experience with Kubernetes, Docker, Terraform, Ansible, and Infrastructure as Code.* Strong understanding of distributed systems, microservices, networking, databases, and cloud architecture.* Hands-on experience with observability platforms such as Prometheus, Grafana, ELK, Splunk, or OpenTelemetry.* Strong understanding of SLI, SLO, SLA, error budgets, availability, latency, reliability, and performance metrics.* Experience designing and operating highly available and fault-tolerant systems.* Strong experience in incident management, troubleshooting, root-cause analysis, and production support.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

TymblHub • Chennai District

On-site
INR 1,500,000 - 2,100,000
Site Reliability Engineer Lead Position
Site Reliability Engineer Lead Position

TymblHub • Hyderabad

On-site
INR 4,500,000 - 7,000,000
Site Reliability Engineer - Lead
Site Reliability Engineer - Lead

TymblHub • Hyderabad

On-site
INR 2,600,000 - 4,000,000
Site Reliability Engineer
Site Reliability Engineer

Qualys • Pune District

Hybrid
INR 1,200,000 - 1,800,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Hilabs • Pune District

On-site
INR 1,500,000 - 2,500,000
Lead Cloud Engineer / Lead Site Reliability Engineer (SRE)
Lead Cloud Engineer / Lead Site Reliability Engineer (SRE)

ITOrizon • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Competitive compensation
Collaborative work culture
Career growth opportunities
Senior Site Reliability Engineer - Cloud Infrastructure
Senior Site Reliability Engineer - Cloud Infrastructure

WITS Innovation Lab • Chandigarh

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

Acesoft Labs • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

Questhiring • Gurugram District

On-site
INR 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

BuildxPartners • Bengaluru Urban

Hybrid
INR 2,400,000 - 4,200,000