Senior Site Reliability Engineer

EPAM Systems

Colombia

On-site

COP 100,000,000 - 180,000,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
Global career opportunities
Paid time off and sick leave
Upskilling and certifications
LinkedIn Learning access

Job summary

EPAM Systems seeks a Senior Site Reliability Engineer to strengthen critical infrastructure reliability and accelerate delivery through robust DevOps practices. You will design reliability strategies, build CI/CD pipelines with GitLab, automate with Python, and operate AWS/Azure cloud environments to meet availability goals.

You will harden practices across networking, security, IAM, and compute, coordinate on-call incident response, and implement Kubernetes deployment patterns to improve

Qualifications

  • 3+ years of site reliability engineering experience in cloud environments.
  • Proven leadership ability to influence reliability practices across teams.
  • Enterprise-scale release management experience supporting frequent deployments.
  • Strong cloud platform knowledge across Amazon Web Services and Microsoft Azure.
  • Advanced Python programming skills for automation and tooling.
  • Solid Kubernetes skills using clusters as a developer.
  • Deep CI/CD and source control knowledge with GitLab or similar DevSecOps platforms.
  • Strong infrastructure fundamentals across networking, compute, security, IAM, and configuration automation.
  • Strong analytical skills for complex problem solving under pressure.
  • Upper-Intermediate English proficiency (B2).
  • Reliable on-call readiness to assess and resolve business-critical issues.

Responsibilities

  • Design reliability strategies for critical infrastructure and services
  • Build and maintain CI/CD pipelines and release workflows using GitLab
  • Automate infrastructure operations and tooling with Python
  • Operate cloud environments across AWS and Azure to meet availability goals
  • Harden and standardize infrastructure practices across networking, security, IAM, and compute
  • Coordinate incident response during on-call shifts and restore service quickly
  • Implement Kubernetes-based deployment and operational patterns to improve stability
  • Analyze reliability signals and root causes to prevent repeat incidents
  • Improve DevOps processes and engineering capabilities to enable faster change
  • Partner with stakeholders to prioritize resilience work over short-term fixes

Skills

SRE experience
Leadership
Release management
AWS
Azure
Python
Kubernetes
CI/CD
GitLab
Networking
Security
IAM
On-call
English (B2)

Tools

GitLab
Kubernetes

Job description

EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.
We are seeking a Senior Site Reliability Engineer to strengthen critical infrastructure reliability and accelerate delivery through robust DevOps practices. You will drive improvements across CI/CD, cloud platforms, and operational readiness while solving high-impact production issues.

Responsibilities
  • Design reliability strategies for critical infrastructure and services
  • Build and maintain CI/CD pipelines and release workflows using GitLab
  • Automate infrastructure operations and tooling with Python
  • Operate cloud environments across AWS and Azure to meet availability goals
  • Harden and standardize infrastructure practices across networking, security, IAM, and compute
  • Coordinate incident response during on-call shifts and restore service quickly
  • Implement Kubernetes-based deployment and operational patterns to improve stability
  • Analyze reliability signals and root causes to prevent repeat incidents
  • Improve DevOps processes and engineering capabilities to enable faster change
  • Partner with stakeholders to prioritize resilience work over short-term fixes
Requirements
  • 3+ years of site reliability engineering experience in cloud environments
  • Proven leadership ability to influence reliability practices across teams
  • Enterprise-scale release management experience supporting frequent deployments
  • Strong cloud platform knowledge across Amazon Web Services and Microsoft Azure
  • Advanced Python programming skills for automation and tooling
  • Solid Kubernetes skills using clusters as a developer
  • Deep CI/CD and source control knowledge with GitLab or similar DevSecOps platforms
  • Strong infrastructure fundamentals across networking, compute, security, IAM, and configuration automation
  • Strong analytical skills for complex problem solving under pressure
  • Upper-Intermediate English proficiency (B2)
  • Reliable on-call readiness to assess and resolve business-critical issues
Nice to have
  • Amazon Web Services expertise, including design patterns for resilient systems
  • Microsoft Azure expertise, including governance and operational best practices
  • AI Architecture experience applied to platform reliability and automation
  • AI Solution Engineering experience for production-grade AI-enabled operations
  • Gen AI Solutions Development experience focused on operational use cases
We offer
  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn

EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer — Cloud & CI/CD Automation
Senior Site Reliability Engineer — Cloud & CI/CD Automation

EPAM Systems • Colombia

On-site
COP 100,000,000 - 180,000,000
Healthcare benefits
Global career opportunities
Paid time off and sick leave
+2
Lead AWS DevOps Engineer
Lead AWS DevOps Engineer

EPAM Systems • Colombia

On-site
COP 323,656,000 - 453,118,000
Healthcare benefits
Employee financial programs
Paid time off and sick leave
+2
Lead Infrastructure Engineer (Automation & Orchestration)
Lead Infrastructure Engineer (Automation & Orchestration)

EPAM Systems • Colombia

On-site
COP 120,000,000 - 180,000,000
Healthcare benefits
Upskilling and certification courses
LinkedIn Learning access
+1
Lead Site Reliability Engineer
Lead Site Reliability Engineer

EPAM Systems • Colombia

Remote
COP 120,000,000 - 220,000,000
Healthcare benefits
Paid time off
LinkedIn Learning access
DevOps Engineer
DevOps Engineer

EPAM Systems • Colombia

On-site
COP 60,000,000 - 120,000,000
Wellness program
100% Payroll Scheme
Legal Benefits
+5
Senior Cloud Engineer (AWS)
Senior Cloud Engineer (AWS)

EPAM Systems • Colombia

On-site
COP 90,000,000 - 180,000,000
Healthcare benefits
Paid time off
Up-skilling and certification courses
+1
Lead Data DevOps
Lead Data DevOps

EPAM Systems • Colombia

On-site
COP 180,000,000 - 240,000,000
Healthcare benefits
Paid time off
Upskilling & certification
+2
Senior Application Support Engineer
Senior Application Support Engineer

EPAM Systems • Colombia

On-site
COP 122,760,000 - 167,400,000
Healthcare benefits
Paid time off
Upskilling & certifications
+4
Senior Node.js Backend Engineer
Senior Node.js Backend Engineer

EPAM Systems • Colombia

On-site
COP 70,000,000 - 110,000,000
Healthcare benefits
Paid time off
LinkedIn Learning
+1
Lead Operational Intelligence Engineer
Lead Operational Intelligence Engineer

EPAM Systems • Colombia

On-site
COP 120,000,000 - 210,000,000
Learning culture
Health coverage
Medical leave coverage
+2