Senior Site Reliability Engineer (SRE)

EPAM Systems

Argentina

Presencial

ARS 1.800.000 - 3.200.000

Jornada completa

hace 25 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura hecha para este puesto de trabajo: un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Healthcare benefits
Paid time off
Upskilling programs
Global career opportunities
LinkedIn Learning access

Descripción de la vacante

EPAM Systems in Argentina is seeking a Senior Site Reliability Engineer (SRE) to bridge software development and operations, automate deployments, and maintain highly available production environments. You will design cloud infrastructure with Terraform, build CI/CD pipelines, monitor systems with Prometheus and Grafana, and respond to incidents to keep services resilient and scalable.

The role requires 3+ years in DevOps or SRE, scripting in Python/Bash/Go/Rust, experience with AWS/Azure/GCP,

Formación

  • 3+ years of experience in systems administration, DevOps, or systems-focused software development.
  • Proficiency in at least one scripting or programming language such as Python, Bash, Go, or Rust.
  • Experience with public cloud providers including AWS, Azure, or GCP, and containerization tools such as Docker and Kubernetes.
  • Understanding of Linux/Unix administration and networking fundamentals including TCP/IP, DNS, and HTTP/SSL/TLS.
  • Familiarity with monitoring and observability tools such as Prometheus, Grafana, and Datadog.
  • Passion for automation, eliminating toil, and building resilient systems that fail gracefully.
  • English proficiency at B2 level or higher.

Responsabilidades

  • Design, build, and maintain cloud infrastructure using IaC practices (Terraform, CloudFormation).
  • Build and optimize CI/CD pipelines to automate deployments and tasks.
  • Design and implement logging, monitoring, and alerting with clear SLOs/SLIs.
  • Respond to production incidents and lead troubleshooting to restore services quickly.
  • Conduct blameless post-mortems to identify root causes and prevent recurrence.
  • Collaborate with developers to optimize system performance and capacity planning.
  • Ensure services scale to handle growth and traffic spikes.

Conocimientos

Python
Bash
Go
Rust
Cloud concepts

Herramientas

Docker
Kubernetes
Terraform
AWS
Azure
GCP
Linux

Descripción del empleo

We are seeking a skilled and proactive Senior Site Reliability Engineer (SRE) to join our engineering team. In this role, you will bridge the gap between software development and systems operations, applying software engineering principles to automate operations, scale infrastructure, and ensure systems remain highly available, resilient, and performant. Your mission is to build, run, and protect the production environments that power our applications, minimizing downtime and helping us deploy software rapidly and safely.

Responsibilities
  • Design, build, and maintain cloud infrastructure using modern IaC practices such as Terraform and CloudFormation
  • Build and optimize CI/CD pipelines to automate software deployments, configuration management, and repetitive operational tasks
  • Design and implement robust logging, monitoring, and alerting systems to establish clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
  • Respond to production incidents and lead troubleshooting efforts to restore services quickly
  • Conduct blameless post-mortems to identify root causes and prevent recurrence
  • Partner with software developers to optimize system performance and plan capacity
  • Ensure services can scale to handle growth and traffic spikes
Requirements
  • 3+ years of experience in systems administration, DevOps, or systems-focused software development
  • Proficiency in at least one scripting or programming language such as Python, Bash, Go, or Rust
  • Experience with public cloud providers including AWS, Azure, or GCP, and containerization tools such as Docker and Kubernetes
  • Understanding of Linux/Unix administration and networking fundamentals including TCP/IP, DNS, and HTTP/SSL/TLS
  • Familiarity with monitoring and observability tools such as Prometheus, Grafana, and Datadog
  • Passion for automation, eliminating toil, and building resilient systems that fail gracefully
  • English proficiency at B2 level or higher
We offer
  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior SRE: Build Resilient Cloud, Automate & Scale
Senior SRE: Build Resilient Cloud, Automate & Scale

EPAM Systems • Argentina

Presencial
ARS 1.800.000 - 3.200.000
Healthcare benefits
Paid time off
Upskilling programs
+2
Senior Site Reliability Engineer/Platform Engineer
Senior Site Reliability Engineer/Platform Engineer

Techunting • Córdoba

Presencial
ARS 3.500.000 - 6.000.000
Site Reliability Engineer
Site Reliability Engineer

AgileEngine • Argentina

Híbrido
ARS 2.000.000 - 4.000.000
Professional growth
Competitive USD-based compensation
Flextime
+1
Senior DevOps Engineer
Senior DevOps Engineer

EPAM Systems • Argentina

Presencial
ARS 1.800.000 - 2.400.000
Healthcare benefits
Global career opportunities
Senior DevOps / Site Reliability Engineer
Senior DevOps / Site Reliability Engineer

Visa Hunt • Argentina

Híbrido
ARS 4.000.000 - 7.000.000
Flexible remote/office options
Competitive salary and compensation
Career growth & mentorship
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Teladoc Health • Argentina

Presencial
ARS 136.170.000 - 196.690.000
Senior Site Reliability Engineer IRC302878
Senior Site Reliability Engineer IRC302878

GlobalLogic • Argentina

Presencial
ARS 1.200.000 - 1.800.000
Exciting Projects
Collaborative Environment
Work-Life Balance
+2
Senior Site Reliability Engineer IRC302878
Senior Site Reliability Engineer IRC302878

GlobalLogic • Buenos Aires

Híbrido
ARS 8.928.000 - 13.392.000
Exciting projects
Collaborative environment
Work-life balance
+2
DevOps Engineer
DevOps Engineer

EPAM Systems • Argentina

Presencial
ARS 1.500.000 - 2.800.000
Healthcare benefits
Paid time off and sick leave
Upskilling and certification courses
+2
Lead DevOps Engineer
Lead DevOps Engineer

EPAM Systems • Argentina

Presencial
ARS 2.000.000 - 5.000.000
Healthcare benefits
Employee financial programs
Paid time off and sick leave
+1