Job Description Site Reliability Engineer (SRE) / Production Support Engineer
Role: Site Reliability Engineer (SRE) / Production Support Engineer
Experience: 5 14 Years
Location: As per business requirements
Employment Type: Full-time
About the Role
We are looking for a highly skilled Site Reliability Engineer (SRE) with strong experience in Production Support, Incident Management, ITIL processes, and Reliability Engineering. The ideal candidate will be responsible for ensuring the availability, performance, scalability, and stability of critical production applications and infrastructure while driving automation and operational excellence.
Key Responsibilities
- Provide L2/L3 production support for business-critical applications and platforms.
- Monitor application and infrastructure health, ensuring high availability and reliability.
- Manage and resolve production incidents, service requests, and problem tickets within defined SLAs.
- Lead troubleshooting, root cause analysis (RCA), and post-incident reviews.
- Drive service reliability improvements through automation and monitoring enhancements.
- Participate in on-call support and major incident management activities.
- Work closely with Development, Infrastructure, Cloud, and Business teams to ensure seamless production operations.
- Implement and maintain observability solutions, alerts, dashboards, and monitoring strategies.
- Support deployment activities and production releases.
- Ensure compliance with ITIL Incident, Problem, Change, and Service Management processes.
- Identify recurring issues and proactively implement preventive solutions.
Required Skills
Site Reliability Engineering
- Strong experience in Site Reliability Engineering (SRE) practices.
- Understanding of SLI, SLO, and SLA concepts.
- Reliability, availability, performance, and capacity management.
- Incident management and problem management.
Production Support
- Extensive experience in Application/Production Support environments.
- Experience managing critical production incidents and service restoration.
- Strong troubleshooting and debugging skills.
- Experience in 24x7 support environments.
ITIL & Service Management
- Good understanding of ITIL Framework.
- Incident, Change, Problem, and Release Management.
- Service Operations and Continual Service Improvement processes.
Cloud & Infrastructure
- Experience with one or more cloud platforms:
- Linux/Unix Administration.
- Container technologies (Docker, Kubernetes).
Monitoring & Observability
- Splunk
- Dynatrace
- AppDynamics
- Prometheus
- Grafana
- ELK Stack
Automation & Scripting
- Shell Scripting
- Python
- PowerShell
- Automation of operational activities
Preferred Skills
- DevOps and CI/CD exposure.
- Jenkins, Git, Maven.
- Kubernetes administration and troubleshooting.
- Database troubleshooting (Oracle, PostgreSQL, SQL Server).
- Kafka, Middleware, Messaging technologies.
- Cloud monitoring and observability tools.
Desired Candidate Profile
- 5-14 years of experience in Production Support and SRE functions.
- Strong expertise in incident management, RCA, and operational excellence.
- Experience supporting large-scale enterprise applications.
- Excellent stakeholder management and communication skills.
- Ability to work under pressure during critical incidents and outages.
- Proactive mindset towards