Senior SRE Lead — Java Cloud & Observability

EPAM Systems

Chile

Presencial

CLP 18.000.000 - 30.000.000

Jornada completa

hace 12 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Destaca para este puesto — genera un currículum y una carta de presentación adaptados en cuestión de un minuto.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Healthcare benefits
Paid time off & sick leave
Upskilling & certification courses
LinkedIn Learning access

Descripción de la vacante

EPAM Systems is seeking a hands-on Lead Site Reliability Engineer to maintain and enhance a Java services ecosystem, partnering with backend engineering to improve reliability, observability, and on-call practices across critical services.

You will provide on-call support for identity services, troubleshoot complex production issues, and deploy patches in cloud infrastructure. Strong AWS, DynamoDB, ElastiCache and Gradle experience are essential.

Formación

  • 5+ years in Site Reliability Engineering or DevOps for distributed systems.
  • Strong experience with Amazon Web Services in production environments.
  • Strong experience with Amazon DynamoDB and Amazon ElastiCache operations.
  • Observability and troubleshooting in distributed systems using logs and telemetry.
  • Hands-on experience with Git-based workflows.
  • Hands-on experience with Gradle in Java service environments.
  • Leadership to guide reliability improvements and support operational decision-making.
  • Incident response skills to communicate operational issues clearly and concisely in writing.
  • Fast learning ability to absorb information quickly and apply it during on-call support.
  • SLO management to track, evaluate, and improve reliability through repeatable processes.
  • English proficiency: B2 (Upper-Intermediate).

Responsabilidades

  • Provide on-call support for Java backend identity services during business hours.
  • Troubleshoot complex production issues using logs and telemetry to identify root causes.
  • Prepare and deploy patches to address issues in cloud infrastructure.
  • Implement reliability improvements for key identity services through practical code and configuration changes.
  • Build and refine metrics and dashboards to enable rapid assessment of platform health.
  • Monitor SLOs across backend services and drive remediation when error rates increase.
  • Create and improve runbooks to standardize operational responses across services.

Conocimientos

SRE/DevOps experience
AWS production
Observability/troubleshooting
Leadership
Incident response
SLO management
English proficiency

Herramientas

Git workflows
Gradle
DynamoDB
ElastiCache
Kubernetes
Terraform
Grafana
New Relic
Kafka

Descripción del empleo

EPAM Systems is seeking a hands-on Lead Site Reliability Engineer to maintain and enhance a Java services ecosystem, partnering with backend engineering to improve reliability, observability, and on-call practices across critical services.

You will provide on-call support for identity services, troubleshoot complex production issues, and deploy patches in cloud infrastructure. Strong AWS, DynamoDB, ElastiCache and Gradle experience are essential.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Lead Elastic Observability Platform Engineer
Lead Elastic Observability Platform Engineer

EPAM Systems • Chile

Presencial
CLP 54.000.000 - 90.000.000
Health insurance
Lunch allowance
Internet & electricity allowance
+5
Remote Senior SRE — Scale, Security & Uptime
Remote Senior SRE — Scale, Security & Uptime

Imachines • Chile

A distancia
CLP 115.385.000 - 173.077.000
Fully remote
Global team
Ship early, ship often
+2
Lead AWS Platform Engineer: Cloud Services & Automation
Lead AWS Platform Engineer: Cloud Services & Automation

EPAM Systems • Chile

Presencial
CLP 60.000.000 - 95.000.000
Healthcare benefits
Global career opportunities
Training & certification programs
+1
Lead Data Engineer (Java + AWS) | Scale Pipelines & Mentorship
Lead Data Engineer (Java + AWS) | Scale Pipelines & Mentorship

EPAM Systems • Chile

Presencial
CLP 6.000.000 - 10.000.000
Improved medical coverage
Lunch Allowance
Internet and electricity allowance
+4
Lead AWS DevOps Engineer
Lead AWS DevOps Engineer

EPAM Systems • Chile

Presencial
CLP 60.000.000 - 95.000.000
Healthcare benefits
Global career opportunities
Training & certification programs
+1
Senior Full-Stack JavaScript Engineer
Senior Full-Stack JavaScript Engineer

EPAM Systems • Chile

Presencial
CLP 88.063.000 - 127.202.000
Healthcare benefits
Paid time off
Upskilling and certification
Hybrid/Remote SRE for Global E-commerce
Hybrid/Remote SRE for Global E-commerce

APPLY • Santiago

Híbrido
CLP 37.106.000 - 64.935.000
Flexible work arrangements
Generous vacation policy
AI & Strategic Upskilling
+2
Remote Hybrid SRE for Global E‑commerce Platforms
Remote Hybrid SRE for Global E‑commerce Platforms

Applydigital • Santiago

Híbrido
CLP 56.075.000 - 84.112.000
Generous vacation policy
Flexible work arrangements
AI upskilling & training budgets
+1
Senior AWS DevOps & SRE | Remote & Flexible Hours
Senior AWS DevOps & SRE | Remote & Flexible Hours

Bluelightconsulting • Antofagasta

A distancia
CLP 28.000.000 - 46.000.000
Work remotely
Flexible working hours
Generous paid time off
+2
Senior DevOps Engineer — Remote, High-Impact SRE
Senior DevOps Engineer — Remote, High-Impact SRE

Bluelight • Concepcion

A distancia
CLP 66.960.000 - 111.600.000
Competitive salary & bonuses
Generous paid time off
Flexible working hours
+2