Senior Site Reliability Engineer

EPAM Systems

Argentina

On-site

ARS 136,430,000 - 181,906,000

Full time

5 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

International projects with top brands
Global teams of diverse peers
Healthcare benefits
Employee financial programs
Paid time off and sick leave
Upskilling & certification courses
LinkedIn Learning access
Global career opportunities
Volunteer and community involvement

Job summary

EPAM Systems seeks a hands-on Senior Site Reliability Engineer to maintain and improve a Java services ecosystem, working with an SRE peer and backend team to boost reliability and observability.

You will contribute to on-call support, build robust runbooks, and implement monitoring with dashboards, SLOs, and incident reporting to reduce downtime and accelerate recovery.

Qualifications

  • 3+ years of Site Reliability Engineering or DevOps experience for distributed systems.
  • On-call support for production services during business hours.
  • Proven experience with Amazon Web Services in production environments.
  • Hands-on with Amazon DynamoDB and Amazon ElastiCache.
  • Strong Git skills for operational and reliability changes.
  • Solid Gradle knowledge for Java-based services.
  • Clear written communication for incident reports.
  • Upper-Intermediate English (B2).

Responsibilities

  • Provide on-call support for Java backend identity services during business hours.
  • Troubleshoot production issues using logs and telemetry and drive root-cause resolution.
  • Prepare and deploy patches to address issues in cloud infrastructure.
  • Improve service reliability by implementing changes that reduce errors and instability.
  • Build and refine metrics and dashboards to surface platform health and behavior.
  • Monitor SLOs and propose code changes to improve attainment.
  • Create and improve runbooks to standardize operations and reduce recovery time.
  • Communicate incidents and risks clearly in writing during live response.
  • Collaborate with engineers to align operational practices with service ownership.

Skills

AWS
DynamoDB
ElastiCache
Git
Gradle
On-call
Incident response
English (B2)

Tools

Kubernetes
Terraform
Grafana
New Relic
Apache Kafka

Job description

We are looking for a hands-on Senior Site Reliability Engineer to help maintain, enhance, and support a Java services ecosystem in close collaboration with an SRE peer and a backend engineering team. You will strengthen reliability, observability, and operational readiness while participating in on-call support.

Responsibilities
  • Provide on-call support for Java backend identity services during business hours
  • Troubleshoot complex production issues using logs and telemetry and drive root-cause resolution
  • Prepare and deploy patches to address issues in cloud infrastructure
  • Improve service reliability by implementing practical changes that reduce errors and instability
  • Build and refine metrics and dashboards to surface platform health and service behavior
  • Monitor SLOs and propose code changes that improve SLO attainment as issues arise
  • Create and improve runbooks to standardize operational response and reduce time to recovery
  • Communicate incidents and operational risks clearly in writing during live response
  • Collaborate closely with engineers to align operational practices with service ownership
Requirements
  • 3+ years of Site Reliability Engineering or DevOps experience supporting distributed systems
  • Strong on-call support experience for production services and incident response during business hours
  • Proven experience with Amazon Web Services in production environments
  • Hands-on experience with Amazon DynamoDB and Amazon ElastiCache
  • Strong Git skills for collaborating on operational and reliability code changes
  • Solid Gradle knowledge for building and maintaining Java-based services
  • Strong troubleshooting skills using logs and telemetry to identify root causes
  • Clear written communication skills for documenting and reporting operational issues during incidents
  • Proactive learning mindset to absorb complex information quickly and apply it under pressure
  • Upper-Intermediate English proficiency (B2)
Nice to have
  • Kubernetes
  • Terraform
  • Grafana
  • New Relic
  • Apache Kafka
We offer
  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

EPAM Systems • Argentina

On-site
ARS 1,400,000 - 2,000,000
Healthcare benefits
Paid time off
Upskilling programs
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

N-iX • Argentina

On-site
ARS 135,892,000 - 196,289,000
Flexible work format
Education reimbursement
Mentorship program
+2
Lead Site Reliability Engineer for Java Identity Services
Lead Site Reliability Engineer for Java Identity Services

EPAM Systems • Argentina

On-site
ARS 1,400,000 - 2,000,000
Healthcare benefits
Paid time off
Upskilling programs
+1
Senior Java Developer
Senior Java Developer

EPAM Systems • Argentina

Hybrid
ARS 1,200,000 - 2,400,000
International projects with top brands
Global teams of skilled peers
Healthcare benefits
+8
Senior Site Reliability Engineer IRC304301
Senior Site Reliability Engineer IRC304301

GlobalLogic • Argentina

Hybrid
ARS 1,800,000 - 3,000,000
Competitive salary
Family medical insurance
Extended paternity leave
+2
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Buenos Aires

On-site
ARS 3,000,000 - 6,000,000
Senior Site Reliability Engineer IRC304301
Senior Site Reliability Engineer IRC304301

t2s - Group International . your partner in executive search • Argentina

On-site
ARS 2,400,000 - 4,200,000
Senior Cloud Engineer (AWS)
Senior Cloud Engineer (AWS)

EPAM Systems • Argentina

On-site
ARS 2,000,000 - 4,500,000
Healthcare benefits
Unlimited LinkedIn Learning access
Paid time off and sick leave
+3
Site Reliability Engineer - Senior Associate (Troubleshooting & Python)
Site Reliability Engineer - Senior Associate (Troubleshooting & Python)

JPMorgan Chase & Co. • Buenos Aires

On-site
ARS 900,000 - 1,500,000
Site Reliability Engineer
Site Reliability Engineer

Strategic Staffing Solutions • Argentina

Remote
ARS 1,200,000 - 2,500,000
Full-time employment
Remote work 100%
Competitive salary in ARS
+1