Job Title
Lead Site Reliability Engineer
Job Details
- Business Unit: Tech & Digital
- Team: Ent Factory-Channels, Mobility, Payments
- Reports to: SRE Manager
- Location: Mumbai, Chennai, Gurgaon & Bangalore
- Role Type: Non-Supervisory
- No of direct reportees: 0
- Travel Required: No
- Job Band Range: D1/D2
- JD Created date: 28th Jan 2023
Job Purpose
Analyzing, troubleshooting, and designing vital services, platforms, and infrastructure on GCP with a focus on reliability, scalability, resilience, security, and performance.
Lead and Mentor a team of SRE engineers.
Job Responsibilities
- Execute reliability initiatives for the team and organization.
- Mentor and lead a team of SRE engineers.
- Help build a Site Reliability Engineering culture by sharing best practices, approaches, documentation, and code with other engineering teams.
- Solid understanding of observability tools and ability to express reliability metrics via observability.
- Define KPI in the form of RPO/RTO/SLI/SLO/Error Budget.
- Apply automation and software to any manually performed tasks or system parts.
- Troubleshoot complicated, cross-platform issues handling OS, Networking, Database in a cloud-based SaaS environment and manage live production incidents.
- Monitor application performance, take steps to improve overall application performance and stability, and follow through with implementation.
- Conduct system analysis, configuration management, and develop improvements for system software performance, availability, and reliability.
- Design, write, ship, and motivate the creation of software and systems to increase observability, product reliability, and organizational efficiency.
- Maintain and monitor deployment, orchestration of servers, docker containers, databases, and general backend infrastructure.
- Develop Run Books/Standard Operating Procedure for recurring Production issues and work on permanent solutions.
- Perform Incident Analysis regularly to prevent and find long-term solutions for Incidents.
Educational Qualifications
B Tech in Computer Science or related discipline preferred.
Key Skills
- Experience in monitoring and analyzing infrastructure performance using standard performance monitoring tools.
- Demonstrable experience in Containerization (Docker) and orchestration (Kubernetes).
- Experience with Infrastructure As Code (Terraform, Cloud Formation, Ansible).
- Knowledge and proven hands-on experience in large-scale databases and distributed technologies, such as Kafka and Confluent Platform Kafka.
- Basic programming and scripting skills.
- Solid understanding of at least 2 observability technologies.
Experience Required
Total Years of experience: 11-13
Major Stakeholders
Internal: Product Manager from Digital Factory, Business Analyst from BTG team, Incident Management team, Development Team.