Job Title: Lead SRE (Site Reliability Engineer)
Location: Gurgaon (Hybrid)
Employment Type: Full-Time
Role Overview
We are looking for an experienced Lead SRE to drive reliability, scalability, and performance of cloud-based systems. The ideal candidate should have strong hands-on experience in Azure-based SRE environments, along with proven leadership capabilities in managing teams and driving operational excellence.
Key Responsibilities
- Lead and manage a team of SRE/DevOps engineers
- Design and implement highly available, scalable, and reliable systems on Azure
- Drive incident management, root cause analysis (RCA), and system improvements
- Automate infrastructure using Terraform and scripting (Shell/Bash)
- Manage Kubernetes clusters and containerized environments (Docker)
- Ensure system monitoring, alerting, and observability best practices
- Collaborate with cross-functional teams and stakeholders for smooth delivery
- Drive performance optimization, cost optimization, and reliability improvements
- Own team management, mentoring, performance reviews, and appraisal cycles
Required Skills
- Strong experience in Site Reliability Engineering (SRE)
- Experience with Terraform (Infrastructure as Code)
- Strong knowledge of Kubernetes and Docker
- Proficiency in Shell / Bash scripting
- Experience with monitoring, logging, and alerting tools
- Strong understanding of CI/CD pipelines and DevOps practices
Leadership & Management
- Proven experience in leading teams and people management
- Experience in appraisal cycles, performance management, and mentoring
- Strong stakeholder management and communication skills
What We’re Looking For
- Ownership mindset with strong problem-solving skills
- Ability to work in a fast-paced, high-impact environment
- Strong collaboration and leadership abilities