Senior Site Reliability Engineer (SRE)

EPAM Systems

Colombia

Presencial

COP 279.520.000 - 403.752.000

Jornada completa

hace 11 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Consigue una respuesta de este empleador — un currículum y una carta de presentación adaptados exactamente a lo que busca para contratar.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Health insurance
Paid time off
Upskilling / certification courses
Global career opportunities
Employee financial programs

Descripción de la vacante

EPAM Systems is seeking a Senior Site Reliability Engineer to join our engineering team. You will bridge software development and operations, automate deployments, scale infrastructure, and ensure high availability, resilience, and performance of production environments powering our applications.

You will design cloud infrastructure with IaC (Terraform, CloudFormation), build CI/CD pipelines, implement robust monitoring and incident response, and collaborate with software engineers to improve

Formación

  • 3+ years of experience in systems administration, DevOps, or systems-focused software development.
  • Proficiency in scripting or programming languages such as Python, Bash, Go, or Rust.
  • Experience with public cloud providers (AWS, Azure, GCP) and containerization tools (Docker, Kubernetes).
  • Understanding of Linux/Unix administration and networking fundamentals (TCP/IP, DNS, HTTP/SSL/TLS).
  • Monitoring and observability tools such as Prometheus, Grafana, and Datadog.
  • English proficiency at B2 level or higher.

Responsabilidades

  • Design, build, and maintain cloud infrastructure using modern IaC practices such as Terraform and CloudFormation
  • Build and optimize CI/CD pipelines to automate software deployments, configuration management, and repetitive operational tasks
  • Design and implement robust logging, monitoring, and alerting systems to establish clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
  • Respond to production incidents and lead troubleshooting efforts to restore services quickly
  • Conduct blameless post-mortems to identify root causes and prevent recurrence
  • Partner with software developers to optimize system performance and plan capacity
  • Ensure services can scale to handle growth and traffic spikes

Descripción del empleo

EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.

We are seeking a skilled and proactive Senior Site Reliability Engineer (SRE) to join our engineering team. In this role, you will bridge the gap between software development and systems operations, applying software engineering principles to automate operations, scale infrastructure, and ensure systems remain highly available, resilient, and performant. Your mission is to build, run, and protect the production environments that power our applications, minimizing downtime and helping us deploy software rapidly and safely.

Responsibilities
  • Design, build, and maintain cloud infrastructure using modern IaC practices such as Terraform and CloudFormation
  • Build and optimize CI/CD pipelines to automate software deployments, configuration management, and repetitive operational tasks
  • Design and implement robust logging, monitoring, and alerting systems to establish clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
  • Respond to production incidents and lead troubleshooting efforts to restore services quickly
  • Conduct blameless post-mortems to identify root causes and prevent recurrence
  • Partner with software developers to optimize system performance and plan capacity
  • Ensure services can scale to handle growth and traffic spikes
Requirements
  • 3+ years of experience in systems administration, DevOps, or systems-focused software development
  • Proficiency in at least one scripting or programming language such as Python, Bash, Go, or Rust
  • Experience with public cloud providers including AWS, Azure, or GCP, and containerization tools such as Docker and Kubernetes
  • Understanding of Linux/Unix administration and networking fundamentals including TCP/IP, DNS, and HTTP/SSL/TLS
  • Familiarity with monitoring and observability tools such as Prometheus, Grafana, and Datadog
  • Passion for automation, eliminating toil, and building resilient systems that fail gracefully
  • English proficiency at B2 level or higher
We offer
  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn

EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior DevOps Engineer
Senior DevOps Engineer

EPAM Systems • Colombia

Presencial
COP 120.000.000 - 180.000.000
Healthcare benefits
Paid time off
Upskilling and certification
+2
DevOps Engineer
DevOps Engineer

EPAM Systems • Colombia

Híbrido
COP 60.000.000 - 120.000.000
Wellness program
100% Payroll Scheme
Legal Benefits
+5
Senior Platform Engineer
Senior Platform Engineer

EPAM Systems • Colombia

Presencial
COP 120.000.000 - 180.000.000
Healthcare benefits
Employee financial programs
Paid time off and sick leave
+2
Senior SRE: Cloud Infra, CI/CD & Resilience
Senior SRE: Cloud Infra, CI/CD & Resilience

EPAM Systems • Colombia

Presencial
COP 279.520.000 - 403.752.000
Health insurance
Paid time off
Upskilling / certification courses
+2
Lead DevOps Engineer
Lead DevOps Engineer

EPAM Systems • Colombia

Presencial
COP 156.240.000 - 267.840.000
Healthcare benefits
Paid time off
Learning & development resources
+1
Lead Operational Intelligence Engineer
Lead Operational Intelligence Engineer

EPAM Systems • Colombia

Presencial
COP 120.000.000 - 210.000.000
Learning culture
Health coverage
Medical leave coverage
+2
Senior DevOps Cloud Engineer (Azure)
Senior DevOps Cloud Engineer (Azure)

EPAM Systems • Colombia

Presencial
COP 240.000.000 - 360.000.000
Healthcare benefits
Learning & development
Global career opportunities
+2
Chief Forward Deployed Engineer
Chief Forward Deployed Engineer

EPAM Systems • Colombia

Presencial
COP 186.347.000 - 372.694.000
Healthcare benefits
Paid time off
Global career opportunities
+1
Site Reliability Engineer
Site Reliability Engineer

Intraway • Bogotá

Presencial
COP 96.000.000 - 140.000.000
Unlimited PTO
Training access
English classes
+2
Lead Platform Engineer
Lead Platform Engineer

EPAM Systems • Colombia

Presencial
COP 120.000.000 - 180.000.000
Healthcare benefits
Career development programs
Paid time off