Lead Site Reliability Engineer

EPAM Systems

Colombia

Remote

COP 120,000,000 - 220,000,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
Paid time off
LinkedIn Learning access

Job summary

EPAM Systems, Inc. is seeking a hands-on Lead Site Reliability Engineer to maintain, enhance, and support a Java services ecosystem while partnering closely with a backend engineering team.

You will strengthen reliability, observability, and on-call practices across critical services. You will lead on-call support for Java identity services, troubleshoot production issues using logs and telemetry, deploy patches to cloud infrastructure, and implement reliability improvements with code and

Qualifications

  • 5+ years of experience in Site Reliability Engineering or DevOps for distributed systems.

Responsibilities

  • Provide on-call support for Java backend identity services during business hours.

Skills

Site Reliability Engineering
DevOps
AWS
DynamoDB
ElastiCache
Observability
Troubleshooting
Git workflows
Gradle
Leadership
Incident response
SLO management
English Proficiency

Tools

Kubernetes
Terraform
Grafana
New Relic
Apache Kafka

Job description

We are looking for a hands-on Lead Site Reliability Engineer to maintain, enhance, and support a Java services ecosystem while partnering closely with a backend engineering team. You will strengthen reliability, observability, and on-call practices across critical services.ResponsibilitiesProvide on-call support for Java backend identity services during business hoursTroubleshoot complex production issues using logs and telemetry to identify root causesPrepare and deploy patches to address issues in cloud infrastructureImplement reliability improvements for key identity services through practical code and configuration changesBuild and refine metrics and dashboards to enable rapid assessment of platform healthMonitor SLOs across backend services and drive remediation when error rates increaseCreate and improve runbooks to standardize operational responses across servicesRequirements5+ years of experience in Site Reliability Engineering or DevOps for distributed systemsStrong experience with Amazon Web Services in production environmentsStrong experience with Amazon DynamoDB and Amazon ElastiCache operationsProven experience with observability and troubleshooting in distributed systems using logs and telemetryHands-on experience with Git-based workflowsHands-on experience with Gradle in Java service environmentsLeadership skills to guide reliability improvements and support operational decision-makingIncident response skills to communicate operational issues clearly and concisely in writingFast learning ability to absorb information quickly and apply it during on-call supportSLO management skills to track, evaluate, and improve reliability through repeatable processesEnglish proficiency: B2 (Upper-Intermediate)Nice to haveKubernetesTerraformGrafanaApache KafkaNew RelicWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedInEPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

EPAM Systems • Colombia

On-site
COP 100,000,000 - 180,000,000
Healthcare benefits
Global career opportunities
Paid time off and sick leave
+2
Senior Full-Stack JavaScript Engineer
Senior Full-Stack JavaScript Engineer

EPAM Systems • Colombia

Remote
COP 120,000,000 - 180,000,000
Healthcare benefits
Paid time off
Upskilling and certification courses
+1
Lead Infrastructure Engineer
Lead Infrastructure Engineer

EPAM Systems • Colombia

Remote
COP 110,000,000 - 170,000,000
Healthcare benefits
Paid time off
Upskilling courses
+5
Senior JavaScript Full-Stack Developer with AWS
Senior JavaScript Full-Stack Developer with AWS

EPAM Systems • Colombia

Remote
COP 120,000,000 - 180,000,000
Healthcare benefits
Paid time off
Lead DevOps Engineer
Lead DevOps Engineer

EPAM Systems • Colombia

Remote
COP 180,000,000 - 320,000,000
Healthcare benefits
Paid time off
Upskilling and certifications
+2
Senior SRE Lead: Reliability, On-Call & Observability
Senior SRE Lead: Reliability, On-Call & Observability

EPAM Systems • Colombia

Remote
COP 120,000,000 - 220,000,000
Healthcare benefits
Paid time off
LinkedIn Learning access
Senior DevOps Engineer
Senior DevOps Engineer

EPAM Systems • Colombia

Remote
COP 120,000,000 - 180,000,000
Healthcare benefits
Paid time off and sick leave
Upskilling and certifications
+1
Java Tech Lead
Java Tech Lead

EPAM Systems • Colombia

Remote
COP 120,000,000 - 210,000,000
Lead AWS DevOps Engineer
Lead AWS DevOps Engineer

EPAM Systems • Colombia

On-site
COP 323,656,000 - 453,118,000
Healthcare benefits
Employee financial programs
Paid time off and sick leave
+2
Senior Data DevOps
Senior Data DevOps

EPAM Systems • Colombia

Remote
COP 110,000,000 - 150,000,000