SRE Engineer - 24/7 Production Reliability & Observability

LTM

Atlanta (GA)

On-site

USD 80,000 - 120,000

Full time

10 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Comprehensive Medical Plan Covering
Short/Long-Term Disability Coverage
401(k) Plan with Company match
Life Insurance
Vacation Time and Paid Holidays
Paid Paternity and Maternity Leave

Job summary

LTIMindtree in the United States is seeking an experienced Production Support Engineer to ensure reliability of AWS-hosted applications. You will handle incident management, triage issues, and collaborate with development teams to restore services quickly.

You will monitor health across CloudWatch, Dynatrace, Quantum Metric, and related tools, build dashboards, and participate in on-call rotations, contributing to robust operational health and continuous improvement.

Qualifications

  • Experience supporting production systems on AWS.
  • Strong incident management and 24/7 support background.
  • Proficient with monitoring/observability tools and dashboards.
  • Troubleshoot across infrastructure, networking, and apps.
  • Familiar with CI/CD and AWS deployment processes.
  • Experience with databases and Unix/Linux environments.

Responsibilities

  • Provide Level 1 and Level 2 production incident support across AWS-hosted apps and infrastructure.
  • Triage incidents by identifying root causes, distinguishing infrastructure issues from application defects, and restoring service within defined SLAs.
  • Escalate code-level defects to development teams with clear diagnostics, supporting logs, and impact assessments.
  • Participate in on-call rotations, major incident bridges, and post-incident reviews.
  • Investigate application defects, configuration issues, and infrastructure anomalies reported through monitoring tools or user incidents.
  • Perform regular health checks across applications, infrastructure, and AWS services.
  • Monitor system health using CloudWatch, Dynatrace, Quantum Metric, and ThousandEyes.
  • Respond proactively to issues related to resource utilization, latency, errors, and availability.
  • Maintain and improve monitoring and observability dashboards.

Skills

Incident management
Production support
AWS infrastructure
Troubleshooting
Networking
Unix/Linux
CI/CD
Monitoring

Tools

CloudWatch
Dynatrace
Quantum Metric
ThousandEyes

Job description

LTIMindtree in the United States is seeking an experienced Production Support Engineer to ensure reliability of AWS-hosted applications. You will handle incident management, triage issues, and collaborate with development teams to restore services quickly.

You will monitor health across CloudWatch, Dynatrace, Quantum Metric, and related tools, build dashboards, and participate in on-call rotations, contributing to robust operational health and continuous improvement.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE Operations Lead — 24x7 Reliability & AI‑Driven Automation
SRE Operations Lead — 24x7 Reliability & AI‑Driven Automation

Tata Consultancy Services • Charlotte (NC)

On-site
USD 70,000 - 120,000
SRE Lead: Production Reliability & Observability Architect
SRE Lead: Production Reliability & Observability Architect

TechDigital Group • Woonsocket (RI)

On-site
USD 140,000 - 190,000
Lead SRE: High-Scale API Reliability & Observability
Lead SRE: High-Scale API Reliability & Observability

LTM • Plano (TX)

On-site
USD 120,000 - 190,000
Comprehensive Medical Plan
401(k) Plan with company match
Vacation & Holidays
+2
Production Reliability Engineer (SRE & Automation)
Production Reliability Engineer (SRE & Automation)

NRnP Technology • Northern (KY)

Hybrid
USD 90,000 - 140,000
Senior Production Reliability Engineer
Senior Production Reliability Engineer

mtb • Buffalo (NY)

On-site
USD 110,000 - 150,000
SRE Lead: Production Reliability & Observability
SRE Lead: Production Reliability & Observability

JPMorganChase • New York (NY)

On-site
USD 140,000 - 190,000
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

On-site
USD 120,000 - 150,000
SRE Lead: Critical Incident Response (Remote)
SRE Lead: Critical Incident Response (Remote)

Liftlab, Inc. • Northern (KY)

Hybrid
USD 140,000 - 190,000
Senior SRE Ops Lead — AI-Driven Reliability & AWS
Senior SRE Ops Lead — AI-Driven Reliability & AWS

Envision Technology Solutions • Charlotte (NC)

On-site
USD 140,000 - 190,000
SRE Engineer
SRE Engineer

Tata Consultancy Services • Englewood Cliffs (NJ)

On-site
USD 110,000 - 125,000