Site Reliability Engineer

Lloyds Technology Centre

Hyderabad

On-site

INR 1,200,000 - 2,400,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Lloyds Technology Centre is seeking a Site Reliability Engineer (SRE) / Production Support Engineer to ensure availability, performance, and stability of critical production applications and infrastructure. The role emphasizes automation, incident management, and strong collaboration with cross-functional teams.

Ideal candidates will have 5–14 years of experience in production support and SRE, with expertise in monitoring, RCA, on-call handling, and ITIL-aligned processes.

Qualifications

  • Strong experience in Site Reliability Engineering (SRE) practices.
  • Expertise in Incident management and RCA.
  • Familiarity with ITIL processes (Incident/Change/Problem/Release).
  • Proficient in automation and scripting to drive reliability.

Responsibilities

  • Provide L2/L3 production support for business-critical applications and platforms.
  • Monitor health of applications and infrastructure to ensure high availability.
  • Manage production incidents, service requests, and problem tickets within SLAs.
  • Lead troubleshooting, RCA, and post-incident reviews.
  • Drive reliability improvements via automation and monitoring enhancements.
  • Participate in on-call and major incident management activities.
  • Collaborate with Development, Infrastructure, Cloud, and Business teams.
  • Implement observability solutions, alerts, dashboards, and monitoring strategies.
  • Support deployment activities and production releases.
  • Identify recurring issues and implement preventive solutions.

Skills

SRE practices
Incident management
ITIL processes
Automation scripting

Tools

Docker
Kubernetes
Splunk
Dynatrace
AppDynamics
Prometheus
Grafana
ELK Stack
Jenkins
Git
Maven

Job description

Job Description Site Reliability Engineer (SRE) / Production Support Engineer

Role: Site Reliability Engineer (SRE) / Production Support Engineer
Experience: 5 14 Years
Location: As per business requirements
Employment Type: Full-time

About the Role

We are looking for a highly skilled Site Reliability Engineer (SRE) with strong experience in Production Support, Incident Management, ITIL processes, and Reliability Engineering. The ideal candidate will be responsible for ensuring the availability, performance, scalability, and stability of critical production applications and infrastructure while driving automation and operational excellence.

Key Responsibilities

  • Provide L2/L3 production support for business-critical applications and platforms.
  • Monitor application and infrastructure health, ensuring high availability and reliability.
  • Manage and resolve production incidents, service requests, and problem tickets within defined SLAs.
  • Lead troubleshooting, root cause analysis (RCA), and post-incident reviews.
  • Drive service reliability improvements through automation and monitoring enhancements.
  • Participate in on-call support and major incident management activities.
  • Work closely with Development, Infrastructure, Cloud, and Business teams to ensure seamless production operations.
  • Implement and maintain observability solutions, alerts, dashboards, and monitoring strategies.
  • Support deployment activities and production releases.
  • Ensure compliance with ITIL Incident, Problem, Change, and Service Management processes.
  • Identify recurring issues and proactively implement preventive solutions.

Required Skills

Site Reliability Engineering

  • Strong experience in Site Reliability Engineering (SRE) practices.
  • Understanding of SLI, SLO, and SLA concepts.
  • Reliability, availability, performance, and capacity management.
  • Incident management and problem management.

Production Support

  • Extensive experience in Application/Production Support environments.
  • Experience managing critical production incidents and service restoration.
  • Strong troubleshooting and debugging skills.
  • Experience in 24x7 support environments.

ITIL & Service Management

  • Good understanding of ITIL Framework.
  • Incident, Change, Problem, and Release Management.
  • Service Operations and Continual Service Improvement processes.

Cloud & Infrastructure

  • Experience with one or more cloud platforms:
    • AWS
    • Azure
    • GCP
  • Linux/Unix Administration.
  • Container technologies (Docker, Kubernetes).

Monitoring & Observability

  • Splunk
  • Dynatrace
  • AppDynamics
  • Prometheus
  • Grafana
  • ELK Stack

Automation & Scripting

  • Shell Scripting
  • Python
  • PowerShell
  • Automation of operational activities

Preferred Skills

  • DevOps and CI/CD exposure.
  • Jenkins, Git, Maven.
  • Kubernetes administration and troubleshooting.
  • Database troubleshooting (Oracle, PostgreSQL, SQL Server).
  • Kafka, Middleware, Messaging technologies.
  • Cloud monitoring and observability tools.

Desired Candidate Profile

  • 5-14 years of experience in Production Support and SRE functions.
  • Strong expertise in incident management, RCA, and operational excellence.
  • Experience supporting large-scale enterprise applications.
  • Excellent stakeholder management and communication skills.
  • Ability to work under pressure during critical incidents and outages.
  • Proactive mindset towards
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Snapmint • Gurugram District

On-site
INR 800,000 - 1,200,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

NOV • Ernakulam

On-site
INR 1,200,000 - 2,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Kameda Infologics Pvt. Ltd. • Thiruvananthapuram

On-site
INR 600,000 - 800,000
Site Reliability Engineer (SRE) – Core IT Infrastructure
Site Reliability Engineer (SRE) – Core IT Infrastructure

TECEZE • Chennai District

On-site
INR 1,000,000 - 2,000,000
Site Reliability Engineer 2
Site Reliability Engineer 2

GreyOrange • Gurugram District

Hybrid
INR 2,600,000 - 4,600,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Synechron • Bengaluru, Hyderabad

Hybrid
INR 4,200,000 - 6,300,000
Software Engineer-DevOps
Software Engineer-DevOps

SMC Squared India • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Sr. Site Reliability Engineer I
Sr. Site Reliability Engineer I

MetLife • Hyderabad

Hybrid
INR 1,500,000 - 2,300,000
Senior Site Reliability Engineer (SRE) / DevOps Engineer
Senior Site Reliability Engineer (SRE) / DevOps Engineer

Umanist Staffing LLC • Maharashtra

On-site
INR 3,500,000 - 5,500,000