Senior Site Reliability Engineer (SRE)

EPAM Systems

Mexico

Hybrid

PHP 900,000 - 1,300,000

Full time

15 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Healthcare benefits
Paid time off and sick leave
Learning & development programs
Global career opportunities
Volunteer opportunities

Job summary

EPAM Systems is seeking a Senior Site Reliability Engineer (SRE) to bridge software development and operations. You will automate infrastructure, scale systems, and ensure high availability across production environments.

You will design cloud infrastructure, CI/CD pipelines, and monitoring, while collaborating with multinational teams to deliver resilient, high-performing solutions. Fluency in English is required.

Qualifications

  • 3+ years of experience in systems administration, DevOps, or SRE-focused software.
  • Proficiency in Python, Bash, Go, or Rust.
  • Experience with AWS, Azure, or GCP and containerization (Docker, Kubernetes).
  • Strong Linux/Unix and networking basics including TCP/IP, DNS, and HTTP/SSL/TLS.
  • English proficiency at B2 level or higher.

Responsibilities

  • Design, build, and maintain cloud infrastructure using IaC like Terraform/CloudFormation.
  • Build and optimize CI/CD pipelines for automated deployments and config management.
  • Implement robust logging, monitoring, and alerting (SLOs/SLIs).
  • Respond to production incidents and lead troubleshooting to restore services.
  • Conduct blameless post-mortems to identify root causes and prevent recurrence.
  • Partner with developers to optimize performance and plan capacity.
  • Ensure services scale to handle growth and traffic spikes.

Skills

Python
Bash
Go
Rust

Tools

Terraform
CloudFormation
Docker
Kubernetes
Prometheus
Grafana
Datadog

Job description

EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.

We are seeking a skilled and proactive Senior Site Reliability Engineer (SRE) to join our engineering team. In this role, you will bridge the gap between software development and systems operations, applying software engineering principles to automate operations, scale infrastructure, and ensure systems remain highly available, resilient, and performant. Your mission is to build, run, and protect the production environments that power our applications, minimizing downtime and helping us deploy software rapidly and safely.

Responsibilities
  • Design, build, and maintain cloud infrastructure using modern IaC practices such as Terraform and CloudFormation
  • Build and optimize CI/CD pipelines to automate software deployments, configuration management, and repetitive operational tasks
  • Design and implement robust logging, monitoring, and alerting systems to establish clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
  • Respond to production incidents and lead troubleshooting efforts to restore services quickly
  • Conduct blameless post-mortems to identify root causes and prevent recurrence
  • Partner with software developers to optimize system performance and plan capacity
  • Ensure services can scale to handle growth and traffic spikes
Requirements
  • 3+ years of experience in systems administration, DevOps, or systems-focused software development
  • Proficiency in at least one scripting or programming language such as Python, Bash, Go, or Rust
  • Experience with public cloud providers including AWS, Azure, or GCP, and containerization tools such as Docker and Kubernetes
  • Understanding of Linux/Unix administration and networking fundamentals including TCP/IP, DNS, and HTTP/SSL/TLS
  • Familiarity with monitoring and observability tools such as Prometheus, Grafana, and Datadog
  • Passion for automation, eliminating toil, and building resilient systems that fail gracefully
  • English proficiency at B2 level or higher
We offer
  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn

EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Site Reliability Engineer
Lead Site Reliability Engineer

EPAM Systems • Mexico

On-site
PHP 7,519,000 - 11,278,000
Healthcare benefits
Paid time off and sick leave
Upskilling and certification courses
+1
Senior SRE – Cloud, CI/CD & Resilient Systems
Senior SRE – Cloud, CI/CD & Resilient Systems

EPAM Systems • Mexico

Hybrid
PHP 900,000 - 1,300,000
Healthcare benefits
Paid time off and sick leave
Learning & development programs
+2
Senior Platform Engineer
Senior Platform Engineer

EPAM Systems • Mexico

On-site
PHP 5,643,000 - 8,150,000
Lead DevOps Engineer
Lead DevOps Engineer

EPAM Systems • Mexico

On-site
PHP 7,524,000 - 11,285,000
Healthcare benefits
Paid time off
Learning & certification opportunities
+2
Site Reliability Engineer
Site Reliability Engineer

PeoplePlusTech Inc. • Metro Manila

Hybrid
PHP 700,000 - 1,100,000
Senior Project Manager
Senior Project Manager

EPAM Systems • Mexico

On-site
PHP 5,643,000 - 8,150,000
Lead DevOps Automation Engineer
Lead DevOps Automation Engineer

EPAM Systems • Mexico

On-site
PHP 1,800,000 - 2,400,000
Healthcare benefits
Paid time off and sick leave
Upskilling, reskilling and certificate
+1
Lead Data DevOps
Lead Data DevOps

EPAM Systems • Mexico

On-site
PHP 7,524,000 - 11,285,000
Healthcare benefits
LinkedIn Learning access
Global career opportunities
+2
Senior Full-Stack Developer
Senior Full-Stack Developer

EPAM Systems • Mexico

On-site
PHP 1,000,000 - 1,900,000
Healthcare benefits
Global career opportunities
Paid time off and sick leave
+2
Senior Full Stack Python Engineer (Integrations)
Senior Full Stack Python Engineer (Integrations)

EPAM Systems • Mexico

Hybrid
PHP 7,524,000 - 11,285,000
Healthcare benefits
LinkedIn Learning access
Global career opportunities
+1