Senior Site Reliability Engineer

Larsen & Toubro Infotech Ltd (LTI)

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Larsen & Toubro Infotech Ltd (LTI) seeks a Senior Site Reliability Engineer to own reliability across AWS and Kubernetes platforms in Bengaluru. You’ll lead incident response, build automation, and scale observability, mentoring engineers as the team grows.

You will design Terraform patterns, drive runbooks, and ensure end-to-end ITSM processes with strong communication to leadership and product teams. AWS certifications are preferred.

Qualifications

  • 8 to 10 years of experience in SRE, DevOps, or cloud infra with production reliability ownership.
  • Hands-on expertise with core AWS services EC2, VPC, IAM, S3, RDS, CloudWatch.
  • Proven troubleshooting of EKS, cluster-level failures, networking, autoscaling, upgrades.
  • Strong Python for building automation and tooling.
  • Extensive Terraform module design and infrastructure patterns at scale.
  • ITSM ownership including Incident and Problem Management.

Responsibilities

  • Lead troubleshooting and resolution of high-severity incidents across AWS and EKS.
  • Design and build Python automation tooling to reduce manual toil.
  • Architect and maintain Terraform infrastructure as code across AWS environments.
  • Drive observability strategy with Grafana dashboards and logs.
  • Perform root-cause analysis using SQL and CloudWatch Logs Insights.
  • Own ITSM processes including incident reviews and remediation plans.
  • Mentor junior and midlevel SREs and collaborate with teams.
  • Design self-healing automation and runbooks to reduce manual recovery time.
  • Monitor across regions for health and failover readiness.

Skills

AWS
EKS
Python
Terraform
SQL
Grafana
ITSM
Incident management
On-call
Communication

Tools

Kubernetes
Terraform
CloudWatch

Job description

Senior Specialist - Architecture Site Reliability Engineer SRE Senior Experience Required 8 to 12 years

About the Role

Were looking for a Senior Site Reliability Engineer to take ownership of reliability performance and operational excellence across our cloud infrastructure and Kubernetes platforms Youll lead the response to complex highseverity incidents raise the bar on automation and observability and mentor other engineers as the team scales

What Youll Do
  • Lead troubleshooting and resolution of complex highseverity incidents across AWS infrastructure and Amazon EKS clusters often serving as incident commander for major incidents
  • Design and build Pythonbased automation frameworks and tooling that eliminate manual toil and improve reliability at scale not just oneoff scripts
  • Architect and maintain Terraformbased infrastructureascode establishing reusable secure and scalable patterns across AWS environments
  • Drive observability strategy design Grafana dashboards and ing frameworks that surface the right signals to the right teams at the right time
  • Use SQL and CloudWatch Logs Insights to perform deep rootcause analysis on complex crossservice incidents
  • Own endtoend ITSM processes Incident Management and Problem Management including root cause analysis postincident reviews and longterm remediation plans
  • Mentor junior and midlevel SREs reviewing their troubleshooting approach automation and incident handling
  • Partner with engineering product and leadership teams to communicate incident impact risk and remediation plans clearly and confidently
  • Design and build selfhealing automation and runbooks that detect known failure patterns and trigger remediation automatically reducing manual intervention and recovery time for recurring incidents
  • Implement and maintain monitoring across multiple regions to ensure consistent visibility into system health latency and failover readiness across all deployment zones
  • Proactively identify potential failure points and performance bottlenecks before they impact production and reduce operational workload by automating recurring manual tasks
  • Participate in and provide seniorlevel escalation support for oncall rotations
What Were Looking For
  • 8 to 10 years of experience in Site Reliability Engineering DevOps or Cloud Infrastructure roles with a track record of owning reliability for productioncritical systems
  • Deep handson expertise with core and advanced AWS services EC2 VPC IAM S3 RDS CloudWatch networking etc
  • Proven expertise troubleshooting complex Amazon EKS issues clusterlevel failures networking autoscaling performance bottlenecks and upgraderelated issues
  • Strong proficiency in Python for building automation frameworks internal tooling and operational systems
  • Extensive experience designing and maintaining Terraform modules and infrastructure patterns at scale
  • Strong command of ITSM frameworks with handson ownership of Incident and Problem Management for highseverity issues
  • Advanced skills querying and analyzing data via SQL and AWS CloudWatch Logs Insights to drive rootcause analysis
  • Proven experience designing Grafana dashboards and ing strategies that scale across multiple teams and services
  • Exceptional verbal and written communication skills able to clearly articulate technical issues risk and remediation plans to engineering leadership and nontechnical stakeholders alike
  • Experience mentoring or leading other engineers and contributing to teamlevel reliability strategy
Mandatory Certifications

AWS Certification required eg AWS Certified Solutions Architect Professional AWS Certified DevOps Engineer Professional or equivalent Certified Kubernetes Administrator CKA or equivalent EKSKubernetes certification required

Soft Skills
  • Calm decisive leadership during highpressure highseverity incidents
  • A strong ownership mindset drives issues to true resolution and follows through on longterm remediation
  • Natural mentor who raises the technical bar for the team
  • Collaborative crossfunctional partner who works effectively with Dev Infra Product and leadership
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE Engineer
Senior SRE Engineer

EPAM Systems India Pvt Ltd • Chennai District

On-site
INR 1,800,000 - 3,200,000
Senior SRE Engineer
Senior SRE Engineer

TymblHub • Chennai District

On-site
INR 2,500,000 - 4,500,000
Site Reliability Engineer
Site Reliability Engineer

PwC Acceleration Center India • Bengaluru

On-site
INR 2,200,000 - 3,800,000
Site Reliability Engineer_ AWS certified
Site Reliability Engineer_ AWS certified

PwC Acceleration Center India • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Site Reliability Engineer
Site Reliability Engineer

Solutions By Text • Bengaluru

On-site
INR 800,000 - 1,200,000
Senior DevSecOps and Site Reliability Engineer
Senior DevSecOps and Site Reliability Engineer

Stryker India Private Limited • Bengaluru

Hybrid
INR 4,000,000 - 6,000,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Lead SRE
Lead SRE

Cvent, Inc. • India

On-site
INR 2,500,000 - 4,500,000
Senior SRE Engineer
Senior SRE Engineer

Epam Systems • Chennai District

On-site
INR 2,500,000 - 4,000,000