# Site Reliability Engineer LeadGood co IndiaPosted on October 2, 2026---Experience5 - 10 yrsSalary (CTC)₹25L - ₹40LJob LocationIndiaVacancy1DesignationSite Reliability Engineer LeadJob TypeNot specified---## Job Description**Role & Responsibilities*** Lead the Site Reliability Engineering function and ensure the availability, reliability, scalability, and performance of critical applications and infrastructure.* Define and implement SRE practices around SLIs, SLOs, SLAs, error budgets, and service reliability.* Design and operate highly available, fault-tolerant, and scalable cloud-native systems.* Lead incident response, troubleshooting, root-cause analysis, and post-incident reviews for critical production issues.* Build and improve monitoring, observability, logging, alerting, and performance-management capabilities.* Drive automation of infrastructure, deployment, operational, and repetitive engineering processes.* Manage and optimize cloud infrastructure across AWS, Azure, GCP, or hybrid environments.* Implement Infrastructure as Code using tools such as Terraform and Ansible.* Manage containerized and orchestration platforms including Docker and Kubernetes.* Establish and improve CI/CD pipelines, deployment strategies, release automation, and rollback mechanisms.* Lead capacity planning, performance optimization, disaster recovery, backup, and business-continuity initiatives.* Identify and reduce operational risks, technical debt, toil, and recurring production issues.* Implement reliability and resilience engineering practices, including failure testing and chaos engineering where appropriate.* Partner with development, infrastructure, security, and product teams to improve application reliability throughout the software lifecycle.* Establish SRE standards, operational runbooks, documentation, and engineering best practices.* Mentor SRE/DevOps engineers and provide technical leadership across reliability initiatives.* Track infrastructure and cloud costs and identify opportunities for performance and cost optimization. **Preferred Candidate Profile*** 5 to 10 years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, or a related field.* Strong hands-on experience managing production systems and large-scale, distributed applications.* Strong proficiency in Linux and scripting/programming languages such as Python, Go, or Bash.* Extensive experience with AWS, Azure, GCP, or hybrid-cloud environments.* Strong experience with Kubernetes, Docker, Terraform, Ansible, and Infrastructure as Code.* Strong understanding of distributed systems, microservices, networking, databases, and cloud architecture.* Hands-on experience with observability platforms such as Prometheus, Grafana, ELK, Splunk, or OpenTelemetry.* Strong understanding of SLI, SLO, SLA, error budgets, availability, latency, reliability, and performance metrics.* Experience designing and operating highly available and fault-tolerant systems.* Strong experience in incident management, troubleshooting, root-cause analysis, and production support.