Lead SRE: Reliability, Observability & On-Call Excellence

EPAM Systems

México

Remote

MXN 900,000 - 1,500,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
Paid time off and sick leave
Upskilling, reskilling and certs
LinkedIn Learning access
Global career opportunities
Volunteer and community involvement
Employee financial programs

Job summary

EPAM Systems, Inc. is seeking a hands-on Lead Site Reliability Engineer to maintain and enhance a Java backend services ecosystem alongside an SRE and backend engineering team. You will strengthen reliability, observability, and incident response practices to ensure robust platform health.

The role focuses on on-call support, troubleshooting, patch deployment in cloud infrastructure, and improving SRE practices with concrete changes. Strong English communication and rapid learning are essential.

Qualifications

  • 5+ years of experience in Site Reliability Engineering or DevOps for distributed systems.
  • 5+ years of experience with Amazon Web Services in production environments.
  • 3+ years of experience with Amazon DynamoDB and Amazon ElastiCache in production.
  • Leadership experience guiding on-call practices and operational improvements.
  • Strong project execution skills delivering reliability and observability enhancements.
  • Strong troubleshooting skills using logs and telemetry to find root causes.
  • Strong communication skills for clear and concise written incident updates.
  • Fast learning ability to absorb information quickly and apply it during on-call.
  • Upper-Intermediate English proficiency (B2)

Responsibilities

  • Provide on-call support for Java backend identity services during business hours.
  • Troubleshoot complex distributed-system issues using logs and telemetry to identify root causes.
  • Deploy patches to address issues in cloud infrastructure.
  • Improve reliability posture for key backend services through tangible changes.
  • Build metrics and dashboards to quickly assess overall platform health.
  • Monitor SLOs across services and submit code changes to improve SLOs as errors occur.
  • Create and refine runbooks for backend services to standardize operations and response.

Skills

On-call leadership
Incident response
Troubleshooting distributed systems
Clear communication
Fast learning

Tools

Amazon Web Services (AWS)
Amazon DynamoDB
Amazon ElastiCache
Kubernetes
Terraform
Apache Kafka
Grafana
New Relic

Job description

EPAM Systems, Inc. is seeking a hands-on Lead Site Reliability Engineer to maintain and enhance a Java backend services ecosystem alongside an SRE and backend engineering team. You will strengthen reliability, observability, and incident response practices to ensure robust platform health.

The role focuses on on-call support, troubleshooting, patch deployment in cloud infrastructure, and improving SRE practices with concrete changes. Strong English communication and rapid learning are essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

EPAM Systems • Mexico

Remote
MXN 900,000 - 1,500,000
Healthcare benefits
Paid time off and sick leave
Upskilling, reskilling and certs
+4
Lead SRE Engineer
Lead SRE Engineer

Cloudsufi • Región Centro

On-site
MXN 900,000 - 1,400,000
Site Reliability Engineer
Site Reliability Engineer

Infojini Inc • Mexico

Remote
MXN 1,200,000 - 1,800,000
Junior SRE: Cloud Reliability & Automation (Azure/K8s)
Junior SRE: Cloud Reliability & Automation (Azure/K8s)

EPAM Systems • Mexico

Remote
MXN 420,000 - 660,000
Lead Systems Reliability Engineer (Enterprise SRE)
Lead Systems Reliability Engineer (Enterprise SRE)

Talent Solutions ManpowerGroup • Ciudad de México

On-site
MXN 900,000 - 1,300,000
Senior SRE Lead: Reliability, Observability & AI Ops
Senior SRE Lead: Reliability, Observability & AI Ops

Cloudsufi • Región Centro

On-site
MXN 900,000 - 1,400,000
Senior Site Reliability Engineer: Build Resilient Systems
Senior Site Reliability Engineer: Build Resilient Systems

Cognizant • Mexico

On-site
MXN 900,000 - 1,500,000
Career growth opportunities
Competitive benefits
Inclusive culture
Remote SRE Engineer - Automate, Observe & Own Reliability
Remote SRE Engineer - Automate, Observe & Own Reliability

AgileEngine • Ciudad de México

Hybrid
MXN 1,049,685 - 1,399,580
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Global SRE: Production Reliability & Automation
Global SRE: Production Reliability & Automation

Fulcrum Digital • Ciudad de México

Remote
MXN 600,000 - 1,000,000
Site Reliability Engineering (SRE) Lead - 2770
Site Reliability Engineering (SRE) Lead - 2770

Xideral • Región Centro

On-site
MXN 700,000 - 900,000
Attractive Salary
Performance bonuses
SGMM Medical insurance