Lead Site Reliability Engineer

EPAM Systems, Inc.

United States

Remote

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
Paid time off
Learning & development
Global career opportunities

Job summary

EPAM Systems, Inc. is seeking a hands-on Lead Site Reliability Engineer to maintain and enhance a Java backend services ecosystem alongside an SRE and backend team.

You will strengthen reliability, observability, and incident response practices across critical services. Your role includes on-call coverage, troubleshooting distributed systems with logs and telemetry, deploying patches in cloud infrastructure, and building dashboards to monitor platform health.

Qualifications

  • 5+ years of Site Reliability Engineering or DevOps for distributed systems.
  • 5+ years of production AWS experience.
  • 3+ years with DynamoDB and ElastiCache in production.
  • Strong on-call leadership and operational improvement experience.
  • Strong troubleshooting using logs and telemetry to identify root causes.
  • Excellent written communication for incident updates.

Responsibilities

  • Provide on-call support for Java backend identity services during business hours.
  • Troubleshoot complex distributed-system issues using logs and telemetry.
  • Deploy patches to address issues in cloud infrastructure.
  • Improve reliability posture of key backend services.
  • Build dashboards and runbooks to monitor health and standardize operations.
  • Monitor SLOs across services and implement code changes to improve them.

Skills

Site Reliability Engineering
On-call leadership
Troubleshooting
Distributed systems
Communication

Tools

AWS
DynamoDB
ElastiCache
Kubernetes
Terraform
Apache Kafka
Grafana
New Relic

Job description

We are looking for a hands-on Lead Site Reliability Engineer to maintain, enhance, and support a Java backend services ecosystem alongside another SRE and a backend engineering team. You will strengthen reliability, observability, and incident response practices.ResponsibilitiesProvide on-call support for Java backend identity services during business hoursTroubleshoot complex distributed-system issues using logs and telemetry to identify root causesDeploy patches to address issues in cloud infrastructureImprove reliability posture for key backend services through tangible changesBuild metrics and dashboards to quickly assess overall platform healthMonitor SLOs across services and submit code changes to improve SLOs as errors occurCreate and refine runbooks for backend services to standardize operations and responseRequirements5+ years of experience in Site Reliability Engineering or DevOps for distributed systems5+ years of experience with Amazon Web Services in production environments3+ years of experience with Amazon DynamoDB and Amazon ElastiCache in productionLeadership experience guiding on-call practices and operational improvementsStrong project execution skills delivering reliability and observability enhancementsStrong troubleshooting skills using logs and telemetry to find root causesStrong communication skills for clear and concise written incident updatesFast learning ability to absorb information quickly and apply it during on-callUpper-Intermediate English proficiency (B2)Nice to haveKubernetesTerraformApache KafkaGrafanaNew RelicWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedIn
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

EPAM Systems, Inc. • United States

Remote
USD 120,000 - 180,000
Healthcare benefits
Paid time off and sick leave
Upskilling and certifications
+1
Senior SRE Lead - Cloud Reliability & Observability
Senior SRE Lead - Cloud Reliability & Observability

EPAM Systems, Inc. • United States

Remote
USD 120,000 - 180,000
Healthcare benefits
Paid time off
Learning & development
+1
Lead DevOps Engineer
Lead DevOps Engineer

EPAM Systems, Inc. • United States

Remote
USD 140,000 - 190,000
Healthcare benefits
Paid time off
Learning & development
+2
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer — Observability & On-Call
Senior Site Reliability Engineer — Observability & On-Call

EPAM Systems, Inc. • United States

Remote
USD 120,000 - 180,000
Healthcare benefits
Paid time off and sick leave
Upskilling and certifications
+1
Senior DevOps Engineer
Senior DevOps Engineer

EPAM Systems, Inc. • United States

Remote
USD 120,000 - 160,000
Healthcare benefits
Employee financial programs
Paid time off and sick leave
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Java SRE Engineer
Java SRE Engineer

Pacer Group • Phoenix (AZ)

On-site
USD 150,000 - 190,000
Medical insurance
Dental insurance
Vision insurance
+1