Lead SRE: High-Scale API Reliability & Observability

LTM

Plano (TX)

On-site

USD 120,000 - 190,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Comprehensive Medical Plan
401(k) Plan with company match
Vacation & Holidays
Life Insurance
Paid parental leave

Job summary

LTIMindtree is seeking an experienced Site Reliability Engineer (SRE) Lead to drive platform reliability, observability, and operational excellence across the API Services ecosystem. This role blends production engineering with reliability leadership and platform security remediation, owning large-scale distributed runtimes.

You will lead reliability efforts, implement SRE best practices, and enhance monitoring and fault detection across global production environments.

Qualifications

  • Strong experience in Site Reliability Engineering and Production Engineering.
  • Hands-on with middleware platforms (MuleSoft/TIBCO) and large distributed runtimes.
  • Deep understanding of system reliability, scalability and high availability design.
  • Experience with observability tools (Dynatrace, Splunk, Prometheus, Grafana).
  • Proficient in CI/CD (Jenkins, Git, Ansible) and scripting (Shell, Python, PowerShell).
  • Experience managing Linux/Unix and Windows production environments.
  • Knowledge of microservices, cloud architectures, and security remediation.

Responsibilities

  • Lead reliability engineering for high-scale API platforms and runtimes.
  • Drive EOL remediation and platform stabilization efforts.
  • Implement SRE best practices, SLIs, SLOs and error budgets.
  • Manage incident response and postmortem processes.
  • Improve observability and proactive fault detection across environments.
  • Support global production with on-call and escalation coverage.

Skills

SRE
Production Eng
Observability
Dynatrace
Splunk
Prometheus
Grafana
Jenkins
Git
Ansible
Shell scripting
Python
PowerShell
Kubernetes
Terraform
Linux/Unix
Windows
MuleSoft
TIBCO

Tools

MuleSoft—TIBCO
APIs

Job description

LTIMindtree is seeking an experienced Site Reliability Engineer (SRE) Lead to drive platform reliability, observability, and operational excellence across the API Services ecosystem. This role blends production engineering with reliability leadership and platform security remediation, owning large-scale distributed runtimes.

You will lead reliability efforts, implement SRE best practices, and enhance monitoring and fault detection across global production environments.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead SRE & Automation Engineering Manager
Lead SRE & Automation Engineering Manager

mtb • Buffalo (NY)

On-site
USD 90,000 - 130,000
SRE Engineer - 24/7 Production Reliability & Observability
SRE Engineer - 24/7 Production Reliability & Observability

LTM • Atlanta (GA)

On-site
USD 80,000 - 120,000
Comprehensive Medical Plan Covering
Short/Long-Term Disability Coverage
401(k) Plan with Company match
+3
Lead SRE Engineer - Reliability & Observability
Lead SRE Engineer - Reliability & Observability

LSEG (London Stock Exchange Group) • Allen (TX)

On-site
USD 170,000 - 210,000
Healthcare
Retirement planning
Volunteer days
+1
Lead Enterprise SRE & Reliability Engineering
Lead Enterprise SRE & Reliability Engineering

Early Warning • Chicago (IL)

Hybrid
USD 194,000 - 237,000
Healthcare coverage
401(k) matching
Paid time off
+2
Lead SRE - Build End-to-End Observability
Lead SRE - Build End-to-End Observability

TrulyHired • Atlanta (GA)

On-site
USD 140,000 - 190,000
Medical insurance
401k match
PTO
Lead Java SRE — Reliability & Observability
Lead Java SRE — Reliability & Observability

KTek Resourcing LLC • Jersey City (NJ)

On-site
USD 140,000 - 170,000
Lead SRE / DevOps Engineer: Observability & Resilience
Lead SRE / DevOps Engineer: Observability & Resilience

Synechron • Dallas (TX)

Hybrid
USD 125,000 - 135,000
Medical insurance
401(k)
Paid maternity leave
+3
SRE Manager: Reliability, Automation & Platform Ops
SRE Manager: Reliability, Automation & Platform Ops

mtb • Buffalo (NY)

Hybrid
USD 150,000 - 230,000
SRE Manager: Lead Reliability & Observability at Scale
SRE Manager: Lead Reliability & Observability at Scale

Iac/interactivecorp • Sacramento (CA)

On-site
USD 150,000 - 210,000
Collaborative work environment
Commitment to carbon‑reduction mission
Flex schedule
+2
Lead SRE: Secure Regulated Cloud & Observability
Lead SRE: Secure Regulated Cloud & Observability

DaParrot Ltd • Northern (KY)

Hybrid
USD 145,000 - 200,000