Senior Site Reliability Engineer

EPAM Systems

United States

Remote

USD 140,000 - 190,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
Paid time off
Learning opportunities
Global teams

Job summary

EPAM Systems, Inc. is seeking a hands-on Senior Site Reliability Engineer to maintain a Java services ecosystem, improve reliability, and participate in on-call support with a backend team.

You will strengthen observability, runbooks, and incident response, troubleshoot production issues from logs and telemetry, and collaborate with engineers to align operational practices with service ownership.

Qualifications

  • 3+ years of Site Reliability Engineering or DevOps experience supporting distributed systems.
  • Strong on-call support experience for production services and incident response during business hours.
  • Hands-on experience with Amazon Web Services in production environments.
  • Experience with DynamoDB and ElastiCache for high‑availability data storage and caching.
  • Proficiency with Git for collaboration on reliability code changes and releases.
  • Gradle knowledge for building and maintaining Java-based services.
  • Excellent written communication for incident reports and post-mortems.
  • Proactive learning mindset to absorb complex information quickly and apply under pressure.

Responsibilities

  • Provide on-call support for Java backend identity services during business hours.
  • Troubleshoot complex production issues using logs and telemetry and drive root-cause resolution.
  • Prepare and deploy patches to address issues in cloud infrastructure.
  • Improve service reliability by implementing practical changes that reduce errors and instability.
  • Build and refine metrics and dashboards to surface platform health and service behavior.
  • Monitor SLOs and propose code changes that improve attainment when issues arise.
  • Create and improve runbooks to standardize operational response and reduce time to recovery.
  • Communicate incidents and operational risks clearly in writing during live response.
  • Collaborate with engineers to align operational practices with service ownership.

Skills

On-call support
Root cause analysis
Incident response
Observability
Communication
English proficiency
SRE
Troubleshooting

Tools

AWS
DynamoDB
ElastiCache
Kubernetes
Terraform
Grafana
New Relic
Apache Kafka

Job description

We are looking for a hands-on Senior Site Reliability Engineer to help maintain, enhance, and support a Java services ecosystem in close collaboration with an SRE peer and a backend engineering team. You will strengthen reliability, observability, and operational readiness while participating in on-call support.ResponsibilitiesProvide on-call support for Java backend identity services during business hoursTroubleshoot complex production issues using logs and telemetry and drive root-cause resolutionPrepare and deploy patches to address issues in cloud infrastructureImprove service reliability by implementing practical changes that reduce errors and instabilityBuild and refine metrics and dashboards to surface platform health and service behaviorMonitor SLOs and propose code changes that improve SLO attainment as issues ariseCreate and improve runbooks to standardize operational response and reduce time to recoveryCommunicate incidents and operational risks clearly in writing during live responseCollaborate closely with engineers to align operational practices with service ownershipRequirements3+ years of Site Reliability Engineering or DevOps experience supporting distributed systemsStrong on-call support experience for production services and incident response during business hoursProven experience with Amazon Web Services in production environmentsHands-on experience with Amazon DynamoDB and Amazon ElastiCacheStrong Git skills for collaborating on operational and reliability code changesSolid Gradle knowledge for building and maintaining Java-based servicesStrong troubleshooting skills using logs and telemetry to identify root causesClear written communication skills for documenting and reporting operational issues during incidentsProactive learning mindset to absorb complex information quickly and apply it under pressureUpper-Intermediate English proficiency (B2)Nice to haveKubernetesTerraformGrafanaNew RelicApache KafkaWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedIn
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

EPAM Systems • United States

Remote
USD 140,000 - 180,000
Kubernetes
Terraform
Apache Kafka
+2
Senior SRE Lead: Reliability, Observability & Cloud
Senior SRE Lead: Reliability, Observability & Cloud

EPAM Systems • United States

Remote
USD 140,000 - 180,000
Kubernetes
Terraform
Apache Kafka
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GoGuardian • El Segundo (CA)

Hybrid
USD 180,000 - 240,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Clearwater Analytics • Boise (ID)

On-site
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000
Site Reliability Engineer
Site Reliability Engineer

Harvey Nash • United States

Remote
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

The ReWork Group • New York (NY)

On-site
USD 120,000 - 160,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Jobgether • United States

On-site
USD 150,000 - 200,000
Competitive salary
Comprehensive healthcare coverage
401(k) plan with company matching
+3
Site Reliability Engineer -- SINDC5717546
Site Reliability Engineer -- SINDC5717546

Compunnel Inc. • Denton (TX)

On-site
USD 120,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

Request Technology, LLC • Chicago (IL)

On-site
USD 150,000 - 155,000