Site Reliability Engineer

EPAM Systems

Argentina

Presencial

ARS 134.672.073 - 179.562.764

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

EPAM Systems is seeking a proactive Site Reliability Engineer to enhance reliability, scalability, and safety of production environments. You will bridge software development and operations through automation, observability, and incident response to reduce downtime and enable fast, safe releases.

Responsibilities include designing cloud infrastructure with IaC, building CI/CD pipelines, and implementing robust monitoring to improve observability.

Formación

  • 2+ years of experience in site reliability engineering, DevOps, or systems administration.
  • Hands-on experience with Infrastructure as Code using Terraform or CloudFormation.
  • Hands-on experience building and improving CI/CD pipelines for automated deployments.
  • Strong troubleshooting and incident response leadership skills in production environments.
  • Solid project skills to coordinate reliability work with software development teams.
  • Proficiency in scripting or programming with Python, Bash, Go, or Rust.
  • Cloud platform experience with AWS, Azure, or GCP.
  • Containerization experience with Docker and Kubernetes.
  • Deep Linux/Unix administration knowledge and networking fundamentals (TCP/IP, DNS, HTTP, SSL/TLS).
  • Strong communication skills with a reliability mindset focused on automation and reducing toil.
  • Advanced English proficiency (C1, Advanced).
  • Nice to have Experience with Prometheus, Grafana, or Datadog.

Responsabilidades

  • Design and maintain cloud infrastructure using Infrastructure as Code practices.
  • Build and optimize CI/CD pipelines to automate deployments and operational workflows.
  • Implement logging, monitoring, and alerting to improve observability and reliability.
  • Define and track Service Level Objectives and Service Level Indicators with clear reporting.
  • Respond to production incidents and drive rapid service restoration.
  • Lead blameless post-mortems to identify root causes and prevent recurrence.
  • Partner with engineers to improve performance, scalability, and capacity planning.
  • Automate repetitive operational tasks to reduce toil and operational risk.
  • Harden production environments to improve resilience and safe change practices.

Conocimientos

CI/CD pipelines
Incident response
Automation
Observability
Python/Bash/Go/Rust
Cloud platforms (AWS/Azure/GCP)
Linux administration
Networking (TCP/IP/DNS/HTTP/SSL)
Communication

Herramientas

Terraform
CloudFormation
Docker
Kubernetes

Descripción del empleo

We are seeking a proactive Site Reliability Engineer to strengthen the reliability, scalability, and safety of production environments. You will bridge software development and operations through automation, observability, and incident response— to help reduce downtime and enable fast, safe releases.

Responsibilities

Design and maintain cloud infrastructure using Infrastructure as Code practices Build and optimize CI/CD pipelines to automate deployments and operational workflows Implement logging, monitoring, and alerting to improve observability and reliability Define and track Service Level Objectives and Service Level Indicators with clear reporting Respond to production incidents and drive rapid service restoration Lead blameless post-mortems to identify root causes and prevent recurrence Partner with engineers to improve performance, scalability, and capacity planning Automate repetitive operational tasks to reduce toil and operational risk Harden production environments to improve resilience and safe change practices

Requirements

2+ years of experience in site reliability engineering, DevOps, or systems administration Hands-on experience with Infrastructure as Code using Terraform or CloudFormation Hands-on experience building and improving CI/CD pipelines for automated deployments Strong troubleshooting and incident response leadership skills in production environments Solid project skills to coordinate reliability work with software development teams Proficiency in scripting or programming with Python, Bash, Go, or Rust Cloud platform experience with AWS, Azure, or GCP Containerization experience with Docker and Kubernetes Deep Linux/Unix administration knowledge and networking fundamentals (TCP/IP, DNS, HTTP, SSL/TLS) Strong communication skills with a reliability mindset focused on automation and reducing toil Advanced English proficiency (C1, Advanced) Nice to have Experience with Prometheus, Grafana, or Datadog

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Site Reliability Engineer/Platform Engineer
Senior Site Reliability Engineer/Platform Engineer

Techunting • Córdoba

Presencial
ARS 89.314.000 - 119.087.000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

EPAM Systems • Argentina

Presencial
ARS 1.800.000 - 3.200.000
Healthcare benefits
Paid time off and sick leave
Upskilling and certification courses
+3
Site Reliability Engineer: Automate, Monitor, and Scale
Site Reliability Engineer: Automate, Monitor, and Scale

EPAM Systems • Argentina

Presencial
ARS 134.672.073 - 179.562.764
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

EPAM Systems • Argentina

Presencial
ARS 133.982.000 - 178.643.000
International projects
Global teams
LinkedIn Learning access
+2
Site Reliability Engineer III
Site Reliability Engineer III

Teladoc Health • Argentina

Híbrido
ARS 1.200.000 - 2.200.000
Site Reliability Engineer: Build Resilient Cloud Infra & CI/CD
Site Reliability Engineer: Build Resilient Cloud Infra & CI/CD

EPAM Systems • Argentina

Presencial
ARS 1.800.000 - 3.200.000
Healthcare benefits
Paid time off and sick leave
Upskilling and certification courses
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Teladoc Health • Argentina

Presencial
ARS 1.800.000 - 2.600.000
Senior SRE: Automate, Scale, and Production Reliability
Senior SRE: Automate, Scale, and Production Reliability

EPAM Systems • Argentina

Presencial
ARS 133.982.000 - 178.643.000
International projects
Global teams
LinkedIn Learning access
+2
Senior Site Reliability Engineer - Remote, Kubernetes
Senior Site Reliability Engineer - Remote, Kubernetes

AgileEngine • Córdoba

Presencial
ARS 3.000.000 - 5.400.000
Growth opportunities
Competitive compensation
Remote work with flexible hours
+3
Senior Site Reliability Engineer - Kubernetes & Observability
Senior Site Reliability Engineer - Kubernetes & Observability

AgileEngine, LLC. • Córdoba

Presencial
ARS 178.643.000 - 267.965.000
Growth without limits
Competitive compensation
Flexibility: 100% remote with flexible
+3