Lead Site Reliability Engineer

EPAM Systems

Argentina

On-site

ARS 1,400,000 - 2,000,000

Full time

17 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
Paid time off
Upskilling programs
Global career opportunities

Job summary

EPAM Systems is seeking a hands-on Lead Site Reliability Engineer to maintain and enhance a Java services ecosystem. You will collaborate with a backend team to strengthen reliability, observability, and on-call practices across critical services.

You will lead on-call support for identity services, troubleshoot production issues via logs and telemetry, and deploy patches for cloud infrastructure. Strong AWS, DynamoDB, ElastiCache experience and leadership are required.

Qualifications

  • 5+ years of experience in Site Reliability Engineering or DevOps for distributed systems.
  • Strong experience with Amazon Web Services in production environments.
  • Strong experience with Amazon DynamoDB and Amazon ElastiCache operations.
  • Proven experience with observability and troubleshooting in distributed systems using logs and telemetry.
  • Hands-on experience with Git-based workflows.
  • Hands-on experience with Gradle in Java service environments.
  • Leadership skills to guide reliability improvements and support operational decision-making.
  • Incident response skills to communicate operational issues clearly and concisely in writing.
  • Fast learning ability to absorb information quickly and apply it during on-call support.
  • SLO management skills to track, evaluate, and improve reliability through repeatable processes.
  • English proficiency: B2 (Upper-Intermediate)

Responsibilities

  • Provide on-call support for Java backend identity services during business hours.
  • Troubleshoot complex production issues using logs and telemetry to identify root causes.
  • Prepare and deploy patches to address issues in cloud infrastructure.
  • Implement reliability improvements for key identity services through practical code and configuration changes.
  • Build and refine metrics and dashboards to enable rapid assessment of platform health.
  • Monitor SLOs across backend services and drive remediation when error rates increase.
  • Create and improve runbooks to standardize operational responses across services.

Skills

SRE/DevOps
Observability
On-call support
Git workflows
Gradle
Incident response
SLO management
English B2

Tools

AWS
DynamoDB
ElastiCache
Gradle
Kubernetes
Terraform
Grafana
New Relic

Job description

We are looking for a hands-on Lead Site Reliability Engineer to maintain, enhance, and support a Java services ecosystem while partnering closely with a backend engineering team. You will strengthen reliability, observability, and on-call practices across critical services.

Responsibilities
  • Provide on-call support for Java backend identity services during business hours
  • Troubleshoot complex production issues using logs and telemetry to identify root causes
  • Prepare and deploy patches to address issues in cloud infrastructure
  • Implement reliability improvements for key identity services through practical code and configuration changes
  • Build and refine metrics and dashboards to enable rapid assessment of platform health
  • Monitor SLOs across backend services and drive remediation when error rates increase
  • Create and improve runbooks to standardize operational responses across services
Requirements
  • 5+ years of experience in Site Reliability Engineering or DevOps for distributed systems
  • Strong experience with Amazon Web Services in production environments
  • Strong experience with Amazon DynamoDB and Amazon ElastiCache operations
  • Proven experience with observability and troubleshooting in distributed systems using logs and telemetry
  • Hands-on experience with Git-based workflows
  • Hands-on experience with Gradle in Java service environments
  • Leadership skills to guide reliability improvements and support operational decision-making
  • Incident response skills to communicate operational issues clearly and concisely in writing
  • Fast learning ability to absorb information quickly and apply it during on-call support
  • SLO management skills to track, evaluate, and improve reliability through repeatable processes
  • English proficiency: B2 (Upper-Intermediate)
Nice to have
  • Kubernetes
  • Terraform
  • Grafana
  • Apache Kafka
  • New Relic
We offer
  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

EPAM Systems • Argentina

On-site
ARS 136,430,000 - 181,906,000
International projects with top brands
Global teams of diverse peers
Healthcare benefits
+6
Lead Site Reliability Engineer for Java Identity Services
Lead Site Reliability Engineer for Java Identity Services

EPAM Systems • Argentina

On-site
ARS 1,400,000 - 2,000,000
Healthcare benefits
Paid time off
Upskilling programs
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

N-iX • Argentina

On-site
ARS 135,892,000 - 196,289,000
Flexible work format
Education reimbursement
Mentorship program
+2
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Buenos Aires

On-site
ARS 3,000,000 - 6,000,000
Senior Site Reliability Engineer IRC304301
Senior Site Reliability Engineer IRC304301

GlobalLogic • Argentina

Hybrid
ARS 1,800,000 - 3,000,000
Competitive salary
Family medical insurance
Extended paternity leave
+2
Senior Site Reliability Engineer IRC304301
Senior Site Reliability Engineer IRC304301

t2s - Group International . your partner in executive search • Argentina

On-site
ARS 2,400,000 - 4,200,000
Senior Java Developer
Senior Java Developer

EPAM Systems • Argentina

Hybrid
ARS 1,200,000 - 2,400,000
International projects with top brands
Global teams of skilled peers
Healthcare benefits
+8
Lead Full-Stack JavaScript Engineer
Lead Full-Stack JavaScript Engineer

EPAM Systems • Argentina

On-site
ARS 106,112,000 - 212,224,000
Healthcare benefits
Upskilling courses
LinkedIn Learning
+2
Tech Lead (.Net)
Tech Lead (.Net)

Skydropx - Frenet • Buenos Aires

Hybrid
ARS 213,233,000 - 289,387,000
Professional growth
USD-based pay
Exciting projects
+1
Site Reliability Engineer - Senior Associate (Troubleshooting & Python)
Site Reliability Engineer - Senior Associate (Troubleshooting & Python)

JPMorgan Chase & Co. • Buenos Aires

On-site
ARS 900,000 - 1,500,000