Stand out for this role — generate a tailored resume and cover letter in about a minute.
Lloyds Banking Group in Hyderabad is seeking a Lead Site Reliability Engineer (F) with 15+ years of experience to drive the reliability strategy for critical products. You will define SLIs/SLOs, lead incident management, and push automation across Java/Python-based tooling.
The role blends software engineering with operations, emphasizing cloud-native practices, observability, and AI-powered operational capabilities to enhance service resilience and customer experience.
End Date
Tuesday 29 September 2026
Hybrid Working
A Lead SRE is accountable for a complex area of the cloud infrastructure resources managing the SLOs through the work of their product team. Advocate for best approach to apply SRE for their technical resources and collaborating with the product teams and the application teams consuming them
15+ Years of experience
A Lead Site Reliability Engineer (SRE) proactively ensures the reliability, availability, scalability, and performance of products deployed in production environments. The role combines software engineering and operational expertise to build, operate, and continuously improve highly resilient digital services.
The Lead SRE is accountable for defining and managing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure services consistently meet business and customer expectations. The role applies software engineering principles to operations, leveraging automation, cloud-native technologies, observability platforms, and AI-powered operational capabilities to improve reliability, reduce operational toil, and enhance customer experience.
The Lead SRE acts as a senior technical leader across one or more products, partnering with engineering, platform, and architecture teams to embed reliability, resilience, observability, and operational excellence into solution design and delivery. The role provides technical leadership during major incidents, problem investigations, and service recovery activities, driving improvements that increase Mean Time To Failure (MTTF) and reduce Mean Time To Restore (MTTR).
Uses deep expertise in reliability engineering, Service Level Objectives (SLOs), Service Level Indicators (SLIs), Error Budgets, incident management, problem management, resilience engineering, and continuous improvement to enhance product reliability and customer experience.
30%Develops automation, operational tooling, APIs, integrations, and self-healing capabilities using Java and/or Python. Applies software engineering principles to reduce operational toil and improve service reliability and efficiency.
25%Designs, operates, and optimises cloud-native platforms and services using Kubernetes/OpenShift, Infrastructure as Code, CI/CD, and modern operational practices to deliver scalable, secure, and highly available solutions.
20%Leverages observability platforms, telemetry analytics, AIOps, and AI-assisted operational capabilities to improve service visibility, incident detection, root cause analysis, predictive insights, and automated remediation.
15%Provides technical leadership, mentoring, and strategic direction across engineering teams. Influences reliability roadmaps, engineering standards, and adoption of SRE best practices whilst fostering a culture of operational excellence.
10%