Principal SRE

Lloyds Bank plc

Hyderabad

Hybrid

INR 4,000,000 - 6,000,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Flexible working options

Job summary

Lloyds Technology Centre in Hyderabad, India, seeks a Lead Site Reliability Engineer to oversee reliability for key products. The role blends software engineering with operations to define SLIs/SLOs, implement error budgets, and drive improvements across incident response and automation.

The Lead SRE will guide cross-functional teams, champion observability, and leverage AI-powered capabilities while mentoring engineers to build resilient services. Flexible hybrid working is offered.

Qualifications

  • 15+ years of experience in SRE or related fields.
  • Experience with Kubernetes/OpenShift and CI/CD.
  • Strong leadership and incident management skills.

Responsibilities

  • Own and drive the reliability strategy for products.
  • Define SLIs, SLOs, and Error Budget policies.
  • Lead major incident management and post-incident reviews.
  • Develop automation and self-healing capabilities in Java/Python.
  • Champion observability and AI-powered operational capabilities.
  • Mentor engineers and collaborate with cross-functional teams.

Skills

SRE
Automation
Cloud
Observability
Mentorship

Tools

Kubernetes/OpenShift
CI/CD

Job description

Lead Site Reliability Engineer (F)

End Date Tuesday 29 September 2026

We Support Flexible Working – Click here for more information on flexible working options

Flexible Working Options Hybrid Working

Job Description Summary

A Lead SRE is accountable for a complex area of the cloud infrastructure resources managing the SLOs through the work of their product team. Advocate for best approach to apply SRE for their technical resources and collaborating with the product teams and the application teams consuming them.

Job Description

A Lead Site Reliability Engineer (SRE) proactively ensures the reliability, availability, scalability, and performance of products deployed in production environments. The role combines software engineering and operational expertise to build, operate, and continuously improve highly resilient digital services. The Lead SRE is accountable for defining and managing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure services consistently meet business and customer expectations. The role applies software engineering principles to operations, leveraging automation, cloud-native technologies, observability platforms, and AI-powered operational capabilities to improve reliability, reduce operational toil, and enhance customer experience. The Lead SRE acts as a senior technical leader across one or more products, partnering with engineering, platform, and architecture teams to embed reliability, resilience, observability, and operational excellence into solution design and delivery.

The role provides technical leadership during major incidents, problem investigations, and service recovery activities, driving improvements that increase Mean Time To Failure (MTTF) and reduce Mean Time To Restore (MTTR).

Role Responsibilities
  • Own and drive the reliability strategy for one or more business-critical products and services.
  • Define, implement, and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budget policies.
  • Lead major incident management, service restoration, and post-incident reviews to drive continual service improvement.
  • Develop and maintain automation, operational tooling, integrations, and self-healing capabilities using Java and/or Python.
  • Drive adoption of cloud-native engineering practices, Infrastructure as Code, CI/CD, and operational automation.
  • Champion observability through monitoring, logging, tracing, telemetry analytics, dashboards, and performance engineering.
  • Leverage AI-powered operational capabilities to improve incident detection, root cause analysis, predictive reliability, and service restoration.
  • Evaluate and recommend new tools, technologies, and engineering practices that improve reliability and operational efficiency.
  • Collaborate with Engineering, Platform, Architecture, and Product teams to embed reliability, resilience, and operability into solution designs.
  • Act as a technical leader and mentor, sharing knowledge and developing engineering capability across teams.
  • Communicate effectively with technical and business stakeholders, influencing reliability investment and prioritisation decisions.
  • Support resilience testing, disaster recovery exercises, operational readiness reviews, and capacity planning activities.
Skill Description Weighting
  • SRE & Service Engineering (30%) – Uses deep expertise in reliability engineering, Service Level Objectives (SLOs), Service Level Indicators (SLIs), Error Budgets, incident management, problem management, resilience engineering, and continuous improvement to enhance product reliability and customer experience.
  • Software Engineering & Automation (Java/Python) (25%) – Develops automation, operational tooling, APIs, integrations, and self-healing capabilities using Java and/or Python. Applies software engineering principles to reduce operational toil and improve service reliability and efficiency.
  • Cloud Platform Engineering (20%) – Designs, operates, and optimises cloud-native platforms and services using Kubernetes/OpenShift, Infrastructure as Code, CI/CD, and modern operational practices to deliver scalable, secure, and highly available solutions.
  • Observability (15%) – Leverages observability platforms, telemetry analytics, AIOps, and AI‑assisted operational capabilities to improve service visibility, incident detection, root cause analysis, predictive insights, and automated remediation.
  • Technical (10%) – Provides technical leadership, mentoring, and strategic direction across engineering teams. Influences reliability roadmaps, engineering standards, and adoption of SRE best practices whilst fostering a culture of operational excellence.

15+ Years of experience Hyderabad Location

We're Lloyds Technology Centre*, a tech and data company located in Hyderabad, India. We're part of Lloyds Banking Group, a leading provider of financial services in the UK and the UK's largest digital bank, with more than 27 million customers. We're changing financial services, and we want you to join us. With market‑leading people practices and great opportunities for career and skills growth, we are committed to creating an exceptional colleague experience that is welcoming for all. Join us! *Lloyds Technology Centre does not offer financial services in India. Should you wish to contact us for any reason, please email us at: indiarecruitment@lloydsbanking.com For more Flexible Working Options please use the free text search, e.g. job sharing, variable hours, to identify relevant matches.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal SRE
Principal SRE

Lloyds Banking Group • Hyderabad

Hybrid
INR 3,000,000 - 5,500,000
Senior SRE
Senior SRE

Lloyds Bank plc • Hyderabad

Hybrid
INR 4,000,000 - 6,000,000
Flexible working
Hybrid working
Site Reliability Engineer
Site Reliability Engineer

LSEG • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

London Stock Exchange Group • Bengaluru

On-site
INR 1,500,000 - 3,000,000
Healthcare
Retirement planning
Paid volunteering days
+1
Senior Cloud Site Reliability Engineer AP
Senior Cloud Site Reliability Engineer AP

Lighthouse • Bengaluru

On-site
INR 1,500,000 - 2,000,000
Annual bonus or incentive program
Diversity and Inclusion initiatives
Flexible working hours
SRE Engineer II
SRE Engineer II

Webhosting • Hyderabad

On-site
INR 1,400,000 - 2,000,000
Lead SRE
Lead SRE

UST • Bengaluru

On-site
INR 3,000,000 - 5,500,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
SRE
SRE

Metlife • Hyderabad

Hybrid
INR 1,500,000 - 2,100,000
Tech & Digital-Lead Site Reliability Engineer
Tech & Digital-Lead Site Reliability Engineer

Hdfc Bank • Bengaluru

On-site
INR 3,500,000 - 5,500,000