Lead Site Reliability Engineer

EPAM Systems

United States

Remote

USD 140,000 - 180,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Kubernetes
Terraform
Apache Kafka
Grafana
New Relic

Job summary

EPAM Systems, Inc. is seeking a hands-on Lead Site Reliability Engineer to maintain and enhance a Java backend services ecosystem in collaboration with an SRE and backend team.

You will improve reliability, observability, and incident response, while leading on-call practices and delivering reliability improvements. The role requires 5+ years in SRE/DevOps, strong AWS experience, DynamoDB/ElastiCache, and leadership in on-call operations.

Qualifications

  • 5+ years of experience in Site Reliability Engineering or DevOps for distributed systems.
  • 5+ years of experience with Amazon Web Services in production environments.
  • Leadership experience guiding on-call practices and operational improvements.
  • Strong project execution skills delivering reliability and observability enhancements.
  • Strong troubleshooting skills using logs and telemetry to find root causes.
  • Strong communication skills for clear and concise written incident updates.

Responsibilities

  • Provide on-call support for Java backend identity services during business hours.
  • Troubleshoot complex distributed-system issues using logs and telemetry to identify root causes.
  • Deploy patches to address issues in cloud infrastructure.
  • Improve reliability posture for key backend services through tangible changes.
  • Build metrics and dashboards to quickly assess overall platform health.
  • Monitor SLOs across services and submit code changes to improve SLOs as errors occur.
  • Create and refine runbooks for backend services to standardize operations and response.

Skills

SRE/DevOps
AWS
Observability
Incident response
On-call leadership
Java backend

Tools

DynamoDB
ElastiCache
Terraform
Kubernetes
Apache Kafka
Grafana
New Relic

Job description

We are looking for a hands-on Lead Site Reliability Engineer to maintain, enhance, and support a Java backend services ecosystem alongside another SRE and a backend engineering team. You will strengthen reliability, observability, and incident response practices.ResponsibilitiesProvide on-call support for Java backend identity services during business hoursTroubleshoot complex distributed-system issues using logs and telemetry to identify root causesDeploy patches to address issues in cloud infrastructureImprove reliability posture for key backend services through tangible changesBuild metrics and dashboards to quickly assess overall platform healthMonitor SLOs across services and submit code changes to improve SLOs as errors occurCreate and refine runbooks for backend services to standardize operations and responseRequirements5+ years of experience in Site Reliability Engineering or DevOps for distributed systems5+ years of experience with Amazon Web Services in production environments3+ years of experience with Amazon DynamoDB and Amazon ElastiCache in productionLeadership experience guiding on-call practices and operational improvementsStrong project execution skills delivering reliability and observability enhancementsStrong troubleshooting skills using logs and telemetry to find root causesStrong communication skills for clear and concise written incident updatesFast learning ability to absorb information quickly and apply it during on-callUpper-Intermediate English proficiency (B2)Nice to haveKubernetesTerraformApache KafkaGrafanaNew RelicWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedIn
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

EPAM Systems • United States

Remote
USD 140,000 - 190,000
Healthcare benefits
Paid time off
Learning opportunities
+1
Senior SRE Lead: Reliability, Observability & Cloud
Senior SRE Lead: Reliability, Observability & Cloud

EPAM Systems • United States

Remote
USD 140,000 - 180,000
Kubernetes
Terraform
Apache Kafka
+2
Lead DevOps Engineer
Lead DevOps Engineer

EPAM Systems • United States

Remote
USD 140,000 - 190,000
Healthcare benefits
Paid time off
Learning & development
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GoGuardian • El Segundo (CA)

Hybrid
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Harvey Nash • United States

Remote
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Axiom Pursuits • San Francisco (CA)

On-site
USD 150,000 - 180,000
Site Reliability Engineer Central Europe
Site Reliability Engineer Central Europe

Zoolatech • Buffalo (NY)

Hybrid
USD 110,000 - 150,000