Site Reliability Engineer (SRE)

EPAM Systems

Colombia

Presencial

COP 60.000.000 - 120.000.000

Jornada completa

Hace 4 días
Sé de los primeros/as/es en solicitar esta vacante

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Healthcare benefits
Global career opportunities
Upskilling and certification courses
LinkedIn Learning access

Descripción de la vacante

EPAM Systems seeks a proactive Site Reliability Engineer to bridge development and operations, design scalable cloud infrastructure, and build robust CI/CD pipelines that automate deployments and configuration management.

You will implement monitoring with Prometheus, Grafana, and Datadog, set SLOs/SLIs, respond to incidents, and collaborate with software teams to ensure resilient, scalable services across global projects.

Formación

  • 2+ years of experience in systems administration, DevOps, or related fields.
  • Proficiency in at least one scripting language such as Python, Bash, Go, or Rust.
  • Experience with AWS, Azure, or GCP and with Docker and Kubernetes.
  • Strong Linux/Unix, networking fundamentals (TCP/IP, DNS, HTTP/SSL/TLS).
  • Advanced English proficiency (C1+).

Responsabilidades

  • Design, build, and maintain cloud infrastructure using modern Infrastructure as Code practices such as Terraform and CloudFormation.
  • Build and optimize CI/CD pipelines to automate software deployments, configuration management, and repetitive operational tasks.
  • Design and implement robust logging, monitoring, and alerting systems using tools such as Prometheus, Grafana, and Datadog.
  • Establish clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs).
  • Respond to production incidents and lead troubleshooting efforts to restore services.
  • Conduct blameless post-mortems to identify root causes and prevent recurrence.
  • Partner with software developers to optimize system performance and plan capacity.
  • Ensure services can scale to handle growth and traffic spikes.

Conocimientos

Automation
Cloud experience
English proficiency

Herramientas

Terraform
CloudFormation
AWS
Azure
GCP
Docker
Kubernetes
Python
Bash
Go
Rust

Descripción del empleo

EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.

Site Reliability Engineer (SRE)

Our engineering team is looking to add a skilled and proactive Site Reliability Engineer (SRE). This position serves as the connective link between software development and systems operations. Software engineering principles will be applied to automate operations, scale infrastructure, and keep systems highly available, resilient, and performant. The core mission involves building, running, and safeguarding the production environments that power our applications, keeping downtime to a minimum while enabling fast, safe software deployment.

Responsibilities
  • Design, build, and maintain cloud infrastructure using modern Infrastructure as Code practices such as Terraform and CloudFormation
  • Build and optimize CI/CD pipelines to automate software deployments, configuration management, and repetitive operational tasks
  • Design and implement robust logging, monitoring, and alerting systems using tools such as Prometheus, Grafana, and Datadog
  • Establish clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
  • Respond to production incidents and lead troubleshooting efforts to restore services
  • Conduct blameless post-mortems to identify root causes and prevent recurrence
  • Partner with software developers to optimize system performance and plan capacity
  • Ensure services can scale to handle growth and traffic spikes
Requirements
  • 2+ years of experience in systems administration, DevOps, or systems-focused software development
  • Proficiency in at least one scripting or programming language such as Python, Bash, Go, or Rust
  • Experience with public cloud providers such as AWS, Azure, or GCP, along with containerization tools such as Docker and Kubernetes
  • Understanding of Linux/Unix administration and networking fundamentals such as TCP/IP, DNS, and HTTP/SSL/TLS
  • Passion for automation, eliminating toil, and building resilient systems that fail gracefully
  • Advanced proficiency in English (C1+)
We offer
  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn

EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Infrastructure Engineer
Senior Infrastructure Engineer

EPAM Systems • Colombia

Presencial
COP 141.199.000 - 188.265.000
Healthcare benefits
Employee financial programs
Paid time off and sick leave
+1
Lead Infrastructure Engineer
Lead Infrastructure Engineer

EPAM Systems • Colombia

Presencial
COP 120.000.000 - 180.000.000
Healthcare benefits
Employee learning programs
Global career opportunities
DevOps Engineer
DevOps Engineer

EPAM Systems • Colombia

Presencial
COP 90.000.000 - 130.000.000
WellHub Plan
100% Payroll
Benefits Package
+3
Site Reliability Engineer: Build Resilient Cloud Infra & CI/CD
Site Reliability Engineer: Build Resilient Cloud Infra & CI/CD

EPAM Systems • Colombia

Presencial
COP 60.000.000 - 120.000.000
Healthcare benefits
Global career opportunities
Upskilling and certification courses
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Metropolitana

Híbrido
COP 142.369.000 - 213.554.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Lead Operational Intelligence Engineer
Lead Operational Intelligence Engineer

EPAM Systems • Colombia

Híbrido
COP 42.000.000 - 68.000.000
Learning Culture
Health Coverage
Stock Option Purchase Plan
+2
Senior Full Stack Developer
Senior Full Stack Developer

EPAM Systems • Colombia

Presencial
COP 60.000.000 - 120.000.000
Healthcare benefits
Global career opportunities
Paid time off & sick leave
+3
Senior .NET Developer
Senior .NET Developer

EPAM Systems • Colombia

Presencial
COP 120.000.000 - 180.000.000
International projects with top brands
Global teams of highly skilled peers
Employee financial programs
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MPS Group LLC • Bogotá

Presencial
COP 200.880.000 - 312.480.000
Senior Java Backend Developer
Senior Java Backend Developer

EPAM Systems • Colombia

Presencial
COP 60.000.000 - 120.000.000
Healthcare benefits
Paid time off
Upskilling & certification