Lead Site Reliability Engineer

EPAM Systems

México

Remote

MXN 900,000 - 1,500,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
Paid time off and sick leave
Upskilling, reskilling and certs
LinkedIn Learning access
Global career opportunities
Volunteer and community involvement
Employee financial programs

Job summary

EPAM Systems, Inc. is seeking a hands-on Lead Site Reliability Engineer to maintain and enhance a Java backend services ecosystem alongside an SRE and backend engineering team. You will strengthen reliability, observability, and incident response practices to ensure robust platform health.

The role focuses on on-call support, troubleshooting, patch deployment in cloud infrastructure, and improving SRE practices with concrete changes. Strong English communication and rapid learning are essential.

Qualifications

  • 5+ years of experience in Site Reliability Engineering or DevOps for distributed systems.
  • 5+ years of experience with Amazon Web Services in production environments.
  • 3+ years of experience with Amazon DynamoDB and Amazon ElastiCache in production.
  • Leadership experience guiding on-call practices and operational improvements.
  • Strong project execution skills delivering reliability and observability enhancements.
  • Strong troubleshooting skills using logs and telemetry to find root causes.
  • Strong communication skills for clear and concise written incident updates.
  • Fast learning ability to absorb information quickly and apply it during on-call.
  • Upper-Intermediate English proficiency (B2)

Responsibilities

  • Provide on-call support for Java backend identity services during business hours.
  • Troubleshoot complex distributed-system issues using logs and telemetry to identify root causes.
  • Deploy patches to address issues in cloud infrastructure.
  • Improve reliability posture for key backend services through tangible changes.
  • Build metrics and dashboards to quickly assess overall platform health.
  • Monitor SLOs across services and submit code changes to improve SLOs as errors occur.
  • Create and refine runbooks for backend services to standardize operations and response.

Skills

On-call leadership
Incident response
Troubleshooting distributed systems
Clear communication
Fast learning

Tools

Amazon Web Services (AWS)
Amazon DynamoDB
Amazon ElastiCache
Kubernetes
Terraform
Apache Kafka
Grafana
New Relic

Job description

We are looking for a hands-on Lead Site Reliability Engineer to maintain, enhance, and support a Java backend services ecosystem alongside another SRE and a backend engineering team. You will strengthen reliability, observability, and incident response practices.ResponsibilitiesProvide on-call support for Java backend identity services during business hoursTroubleshoot complex distributed-system issues using logs and telemetry to identify root causesDeploy patches to address issues in cloud infrastructureImprove reliability posture for key backend services through tangible changesBuild metrics and dashboards to quickly assess overall platform healthMonitor SLOs across services and submit code changes to improve SLOs as errors occurCreate and refine runbooks for backend services to standardize operations and responseRequirements5+ years of experience in Site Reliability Engineering or DevOps for distributed systems5+ years of experience with Amazon Web Services in production environments3+ years of experience with Amazon DynamoDB and Amazon ElastiCache in productionLeadership experience guiding on-call practices and operational improvementsStrong project execution skills delivering reliability and observability enhancementsStrong troubleshooting skills using logs and telemetry to find root causesStrong communication skills for clear and concise written incident updatesFast learning ability to absorb information quickly and apply it during on-callUpper-Intermediate English proficiency (B2)Nice to haveKubernetesTerraformApache KafkaGrafanaNew RelicWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedInEPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Junior Site Reliability Engineer
Junior Site Reliability Engineer

EPAM Systems • Mexico

Remote
MXN 420,000 - 660,000
Lead SRE: Reliability, Observability & On-Call Excellence
Lead SRE: Reliability, Observability & On-Call Excellence

EPAM Systems • Mexico

Remote
MXN 900,000 - 1,500,000
Healthcare benefits
Paid time off and sick leave
Upskilling, reskilling and certs
+4
Senior Full-Stack JavaScript Engineer
Senior Full-Stack JavaScript Engineer

EPAM Systems • Mexico

Remote
MXN 900,000 - 1,200,000
Lead DevOps Engineer
Lead DevOps Engineer

EPAM Systems • Mexico

Remote
MXN 1,000,000 - 1,800,000
International projects
Healthcare benefits
Paid time off and sick leave
+1
Java Tech Lead
Java Tech Lead

EPAM Systems • Mexico

Remote
MXN 1,000,000 - 2,000,000
Senior Data Software Engineer
Senior Data Software Engineer

EPAM Systems • Mexico

Remote
MXN 900,000 - 1,200,000
Lead Full-Stack JavaScript Engineer
Lead Full-Stack JavaScript Engineer

EPAM Systems • Mexico

On-site
MXN 900,000 - 1,600,000
Healthcare benefits
Paid time off and sick leave
Upskilling and certification courses
+2
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Rosarito

On-site
MXN 870,019 - 1,305,028
Professional growth
Competitive compensation
Exciting projects
+1
Lead AWS DevOps Engineer
Lead AWS DevOps Engineer

EPAM Systems • Mexico

On-site
MXN 600,000 - 1,000,000
Healthcare benefits
Global teams
Paid time off and sick leave
+6
Lead Data DevOps
Lead Data DevOps

EPAM Systems • Mexico

Remote
MXN 900,000 - 1,500,000
Healthcare benefits
Paid time off and sick leave
Upskilling & certification courses
+2