Job Title: Site Reliability Engineer
Job Details:
- Business Unit: Tech & Digital
- Team: Ent Factory-Channels, Mobility, Payments
- Reports to: SRE Manager
- Location: Bangalore
- Job Function: Support
- Role Type: Individual contributor
- No of direct reportees: 0
- Travel Required: No
- Job Band Range: E2/E3
- JD Created date: 28th Jan 2023
- JD Updated date: May 8 2024
- Version No: 2.0
Job Purpose
Analyzing, troubleshooting, and designing vital services, platforms, and infrastructure with a focus on reliability, scalability, resilience, security, and performance.
Job Responsibilities
- Help build a Site Reliability Engineering culture by sharing best practices, approaches, documentation, and code with other engineering teams.
- Apply automation and software to tasks or system parts performed manually.
- Troubleshoot complicated, cross-platform issues handling OS, Networking, Database in a cloud-based SaaS environment and manage live production incidents.
- Monitor application performance, take steps to improve overall application performance and stability, and follow through with implementation.
- Conduct system analysis, configuration management, and develop improvements for system software performance, availability, and reliability.
- Design, write, ship, and motivate the creation of software and systems to increase observability, product reliability, and organizational efficiency.
- Maintain and monitor deployment, orchestration, of servers, docker containers, databases, and general backend infrastructure.
- Develop Run Books/Standard Operating Procedure for recurring Production issues, also working on a permanent solve.
- Perform Incident Analysis regularly to prevent and find long-term solutions for Incidents.
Educational Qualifications
B Tech in Computer Science or related discipline preferred.
Key Skills
- Experience in monitoring and analyzing infrastructure performance using standard performance monitoring tools.
- Demonstrable experience in Containerization-Docker and orchestration (Kubernetes).
- Experience with Infrastructure As Code (Terraform, Cloud Formation, Ansible).
- Knowledge and proven hands-on experience in large-scale databases and distributed technologies, such as Kafka and Confluent Platform Kafka.
- Basic programming and scripting skills.
Experience Required
1-3 Total Years of experience.
Major Stakeholders
Internal:
- Product Manager from Digital Factory
- Business Analyst from BTG team
- Incident Management team
- Development Team
The summary has been successfully posted in the Job Description field.
Required Skills
Refer to the Job Description