Senior Team Lead | Site Reliability Engineering | Bengaluru | Engineering

Deloitte & Touche GmbH Wirtschaftsprüfungsgesellschaft

Bengaluru

On-site

INR 4,000,000 - 6,000,000

Full time

18 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Deloitte & Touche GmbH Wirtschaftsprüfungsgesellschaft is seeking a Senior Team Lead in Bengaluru to drive Site Reliability Engineering for critical, distributed systems. You will own end-to-end reliability, scalability and performance while coordinating with cross-functional teams in a 24x7 on-call model.

The role requires deep cloud expertise (AWS), container orchestration (Kubernetes), and strong automation using Terraform, CI/CD pipelines, and observability tooling to reduce MTTR and improve

Qualifications

  • Bachelor's in Engineering.
  • 5–8 years of experience in SRE/DevOps/Cloud Engineering.
  • Hands-on experience managing production-grade systems (24x7).

Responsibilities

  • Own end-to-end production systems reliability, availability, scalability, cost and performance.
  • Participate in 24x7 on-call rotations and handle high-severity incidents.
  • Design, deploy and manage infrastructure on AWS.
  • Deploy and manage containerized workloads using Kubernetes.
  • Implement and manage infrastructure using Terraform.

Skills

AWS
Kubernetes
Terraform
Docker
Jenkins
GitHub
Python
Shell scripting
Dynatrace
Grafana
BigQuery
Pub/Sub
Canary
Blue-Green
Java
Golang
TCP/IP
SLI/SLO/SLA
MTTR/MTTA

Education

Bachelors in Engineering

Tools

Terraform
Jenkins
GitHub
Docker

Job description

Select how often (in days) to receive an alert:

Job Title: Senior Team Lead | Site Reliability Engineering | Bengaluru | Engineering

The team

Engineering helps empower and drive mission-critical solutions whether we need to modernize existing systems or implement new technology products and platforms. Through innovation, we improve financial performance, accelerate new digital businesses and fuel growth. Learn more about Engineering, AI and Data

  • We are looking for a highly skilled Site Reliability Engineer (SRE) to manage and scale mission-critical, production-grade distributed systems running on AWS. The ideal candidate will focus on reliability, automation, observability, and operational excellence while minimizing toil and improving system availability. Maintaining and improving 4 Nines of uptime to 5 Nines with engineering efforts.
  • This role requires deep technical expertise in cloud-native technologies, Kubernetes, infrastructure automation, Linux administration and TCP/IP fundamentals, programming in one language and strong troubleshooting capabilities for distributed systems. The candidate needs to participate in the overall lifecycle management of mission critical banking services with a 24x7 operations mode in an rotational on-call basis. The job requires the candidate to have strong troubleshooting skills in a distributed environment spread across multiple cloud environments. The bare minimum ask would be to maintain high level of agility, learnability and adaptability in different scenarios. An engineer with a zeal to learn fast and having a bias for action would be the best fit for the role.
  • Education-Bachelors in Engineering
  • Reliability & Operations
  • Own end-to-end production systems reliability, availability, scalability, cost and performance.
  • Drive measurable improvements in MTTR, MTTA, and incident response practices using automation and runbook additions and process enhancements.
  • Participate in 24x7 on-call rotations and handle high-severity incidents and document the learnings on ongoing basis.
  • Establish and manage SLI, SLO, SLA, Error Budgets, and operational metrics for mission critical services and partner with engineering teams with full accountability for upholding the SLOs.
  • Partner with the various engineering, operations and cloud management teams to deliver highly reliable service in a timely manner.
  • Design, deploy, and manage infrastructure on AWS
  • Work extensively on:
  • Compute, networking, IAM, Load Balancers, TLS Certs
  • BigQuery, Pub/Sub, cloud logging enhancement, metrics and logs analysis
  • Implement and manage infrastructure using Terraform (Infrastructure as Code).
  • Deploy and manage containerized workloads using Kubernetes.
  • Troubleshoot issues related to:
  • Pods, nodes, networking, storage, services on an ongoing basis
  • Manage deployments using Helm, YAML, and rollout strategies (Canary/Blue-Green).
  • Automation & CI/CD
  • Build and maintain CI/CD pipelines using:
  • Jenkins (pipeline-based, Groovy / Shell / Python scripting)
  • Strong experience in using GitHub as a PowerUser
  • Develop automation using Python and Shell scripting.
  • Reduce operational toil through automation initiatives.
  • Observability & Monitoring
  • Implement and manage monitoring systems using:
  • Dynatrace, Grafana, logs and metrics explorer
  • Work with logs, metrics, and traces for deep observability to identify trends and arrest problems proactively.
  • Define alerting strategies based on system behaviour and SLOs and create runbooks.
  • Work alongside operations teams to identify, fix the production incidents and own the problem resolution.
  • Work with engineering teams to isolate infra and application issues and set up right tooling for debugging production incidents.
  • System & Application Troubleshooting
  • Distributed systems
  • Microservices-based architectures on containerised workloads
  • Java and Golang applications
  • Application issues
  • Infrastructure issues
  • Network-related problems
  • Plan and execute continuous improvement
  • Identify and eliminate repetitive manual tasks.
  • Drive reliability engineering practices and culture. (DRY – Don’t Repeat Yourself)
  • Collaborate with development teams to improve system design and resilience.
  • Experience
  • 5-8 years of relevant and progressive experience in SRE / DevOps / Cloud Engineering
  • Hands‑on experience managing production‑grade systems (24x7 environments)
  • Technical Skills
  • Strong expertise in AWS
  • Kubernetes Engine, VPC, IAM, Load Balancing, LB, Certs, KMS, logs and metrics exploration
  • BigQuery, Pub/Sub
  • Good understanding of cloud architecture and landing zones
  • Infrastructure as Code
  • Strong hands‑on experience with Terraform
  • Ability to write and debug Terraform code from scratch
  • Containers & Orchestration
  • Deep expertise in:
  • Docker
  • Strong troubleshooting experience in Kubernetes environments
  • CI/CD & Automation
  • Hands‑on experience with:
  • Jenkins (pipeline-based CI/CD)
  • GitHub
  • Python (preferred)
  • Shell scripting
  • Experience with automation frameworks and tooling
  • Observability
  • Experience with:
  • Dynatrace / Grafana
  • Log, metrics, and trace‑based monitoring
  • Working knowledge of:
  • Java and/or Golang applications
  • Strong debugging skills across application and infrastructure layers
  • Deep understanding of TCP/IP networking
  • Ability to debug network issues in distributed system
  • Reliability Engineering Skills
  • Solid understanding of:
  • SLI, SLO, SLA, Error Budgets
  • Demonstrable and Proven Experience improving:
  • MTTR, MTTA
  • Experience handling incident management lifecycle
  • Soft Skills
  • Strong analytical and troubleshooting mindset
  • Excellent communication and stakeholder management
  • Ability to work in high‑pressure production environments
  • Ownership‑driven and proactive approach
  • Preferred Qualifications
  • Exposure to banking/financial domain (optional but valuable)
  • Understanding of security and compliance practices
  • Experience with deployment strategies:
  • Canary, Blue‑Green
  • Ideal Candidate Profile in summary would be like :
  • Hands‑on production troubleshooting expert
  • Good at automation + reducing toil
  • Deep understanding of SRE principles
  • Comfortable in 24x7 production environments
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hilabs • Bengaluru

On-site
INR 2,200,000 - 3,500,000
Senior Site Reliability Engineer (SRE) Engineer
Senior Site Reliability Engineer (SRE) Engineer

Umanist Staffing • Pune District

On-site
INR 2,250,000 - 2,750,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Software Engineer
Software Engineer

PwC • Hyderabad, Bengaluru

Hybrid
INR 2,800,000 - 5,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

BayOne Solutions • Hyderabad

On-site
INR 3,000,000 - 6,000,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Site Reliability Engineer (SRE) Engineer
Senior Site Reliability Engineer (SRE) Engineer

Umanist NA • Maharashtra

On-site
INR 1,500,000 - 2,500,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Hilabs • Pune District

On-site
INR 1,500,000 - 2,500,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000