Senior Site Reliability Engineer

EPAM Systems, Inc.

United States

Remote

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
Paid time off and sick leave
Upskilling and certifications
LinkedIn Learning access

Job summary

EPAM Systems, Inc. seeks a hands-on Senior Site Reliability Engineer to maintain, enhance, and support a Java services ecosystem.

The role emphasizes reliability, observability, and operational readiness, with on-call collaboration alongside an SRE peer and backend engineers. Responsibilities include on-call coverage for identity services, root-cause troubleshooting, patch deployment in cloud, and building runbooks and dashboards to surface platform health.

Qualifications

  • 3+ years of Site Reliability Engineering or DevOps experience with distributed systems.
  • Experience with on-call incident response during business hours.
  • Hands-on AWS production experience (EC2, IAM, VPC, RDS).
  • Proven debugging with logs/telemetry and root-cause analysis.

Responsibilities

  • Provide on-call support for Java backend services during business hours.
  • Troubleshoot production issues using logs and telemetry; perform root-cause analysis.
  • Prepare and deploy patches to cloud infrastructure to fix issues.
  • Improve reliability by implementing changes to reduce errors and instability.

Skills

On-call support
AWS
DynamoDB
ElastiCache
Git
Gradle
Troubleshooting
English (B2)

Tools

Kubernetes
Terraform
Grafana
New Relic
Apache Kafka

Job description

We are looking for a hands-on Senior Site Reliability Engineer to help maintain, enhance, and support a Java services ecosystem in close collaboration with an SRE peer and a backend engineering team. You will strengthen reliability, observability, and operational readiness while participating in on-call support.ResponsibilitiesProvide on-call support for Java backend identity services during business hoursTroubleshoot complex production issues using logs and telemetry and drive root-cause resolutionPrepare and deploy patches to address issues in cloud infrastructureImprove service reliability by implementing practical changes that reduce errors and instabilityBuild and refine metrics and dashboards to surface platform health and service behaviorMonitor SLOs and propose code changes that improve SLO attainment as issues ariseCreate and improve runbooks to standardize operational response and reduce time to recoveryCommunicate incidents and operational risks clearly in writing during live responseCollaborate closely with engineers to align operational practices with service ownershipRequirements3+ years of Site Reliability Engineering or DevOps experience supporting distributed systemsStrong on-call support experience for production services and incident response during business hoursProven experience with Amazon Web Services in production environmentsHands-on experience with Amazon DynamoDB and Amazon ElastiCacheStrong Git skills for collaborating on operational and reliability code changesSolid Gradle knowledge for building and maintaining Java-based servicesStrong troubleshooting skills using logs and telemetry to identify root causesClear written communication skills for documenting and reporting operational issues during incidentsProactive learning mindset to absorb complex information quickly and apply it under pressureUpper-Intermediate English proficiency (B2)Nice to haveKubernetesTerraformGrafanaNew RelicApache KafkaWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedIn
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

EPAM Systems, Inc. • United States

Remote
USD 120,000 - 180,000
Healthcare benefits
Paid time off
Learning & development
+1
Senior Site Reliability Engineer — Observability & On-Call
Senior Site Reliability Engineer — Observability & On-Call

EPAM Systems, Inc. • United States

Remote
USD 120,000 - 180,000
Healthcare benefits
Paid time off and sick leave
Upskilling and certifications
+1
Senior DevOps Engineer
Senior DevOps Engineer

EPAM Systems, Inc. • United States

Remote
USD 120,000 - 160,000
Healthcare benefits
Employee financial programs
Paid time off and sick leave
+2
Senior SRE Lead - Cloud Reliability & Observability
Senior SRE Lead - Cloud Reliability & Observability

EPAM Systems, Inc. • United States

Remote
USD 120,000 - 180,000
Healthcare benefits
Paid time off
Learning & development
+1
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

O.C. Tanner • Salt Lake City (UT)

On-site
USD 180,000 - 240,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Harvey Nash • United States

Remote
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

On-site
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3