An application made for this job — a tailored resume and cover letter that speak straight to the posting.
HDFC Bank is seeking a Senior Site Reliability Engineer to analyse, troubleshoot, and design vital services and infrastructure with a focus on reliability, scalability, resilience, security, and performance. You will work on cloud-based SaaS environments, automate tasks, and drive observability across containers and backend systems.
The role requires 8–10 years of experience, strong Linux and networking fundamentals, and expertise in Docker/Kubernetes, Terraform, and CI/CD pipelines.
Senior Site Reliability Engineer
Business Unit: Tech & Digital
Team: DTIT - Enterprise Factory
Reports to: Lead Site Reliability Engineer
Location: Mumbai, Chennai, Gurgaon & Bangalore
Role Type: Individual Contributor
No of direct reportees: NIL
Travel Required: No
Job Band Range: E4
Analysing, troubleshooting, and designing vital services, platforms, and infrastructure while always thinking about reliability, scalability, resilience, security, and performance.
Help build a Site Reliability Engineering culture by sharing best practices, approaches, documentation, and code with other engineering teams.
Apply automation and software to any tasks or parts of the system which are performed manually.
Able to troubleshoot complicated, cross-platform issues handling OS, Networking, Database in a cloud-based SaaS environment and handle live production incidents.
Monitor application performance, take steps to improve overall application performance and stability, and follow through with implementation.
Conduct system analysis, configuration management, and develop improvements for system software performance, availability, and reliability.
Design, write, ship, and motivate the creation of software and systems to increase observability, product reliability, and organizational efficiency.
Maintain and monitor deployment, orchestration, of servers, docker containers, databases, and general backend infrastructure.
Develop Run Books/Standard Operating Procedure for recurring Production issues, also working on a permanent solve.
Perform Incident Analysis on a regular basis with the intention of preventing and finding a long-term solve for Incidents.
B Tech in Computer Science or related discipline preferred.
Experience in monitoring and analyzing infrastructure performance using standard performance monitoring tools.
Demonstrable experience in Containerization-Docker and orchestration (Kubernetes).
Experience with Infrastructure As Code (Terraform, Cloud Formation, Ansible).
Knowledge and proven hands-on experience in large-scale databases and distributed technologies, such as Kafka and Confluent Platform Kafka.
Basic programming and scripting skills.
Total Yrs of experience: 8-10
Internal:
Product Manager from Digital Factory
Business Analyst from BTG team
Incident Management team
Development Team