Lead Site Reliability Engineer (SRE)

EPAM Systems

Mexico

On-site

PHP 5,846,153 - 8,615,384

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

EPAM Systems is seeking a Lead Site Reliability Engineer to ensure cloud platforms are reliable, scalable, and secure through software-driven operations and automation. You will reduce toil, improve availability, and protect production for teams delivering financial services, insurance, and retail solutions.

You will design IaC with Terraform or CloudFormation, build CI/CD pipelines, and implement robust logging and monitoring to define SLOs/SLIs.

Qualifications

  • 5+ years in systems administration or DevOps
  • Proficient in at least one scripting language (Python/Bash/Go/Rust)
  • Experience with cloud providers and containerization
  • Strong knowledge of Linux/Unix networking

Responsibilities

  • Design, build and maintain cloud infrastructure using IaC (Terraform, CloudFormation)
  • Create and optimize CI/CD pipelines for automated deployments
  • Implement logging, monitoring, and SLO/SLI definitions
  • Respond to incidents and lead post-mortems to prevent recurrence
  • Collaborate with developers to optimize performance and scale

Skills

Automation mindset
Toil reduction
English proficiency (B2+)
Collaboration

Tools

Terraform
CloudFormation
AWS
Azure
GCP
Docker
Kubernetes
Prometheus
Grafana
Datadog

Job description

We are seeking a Lead Site Reliability Engineer (SRE) to keep our cloud platforms reliable, scalable, and safe through software-driven operations and automation. You will reduce toil, strengthen availability, and protect production for teams delivering solutions across Financial Services, Insurance, and Retail. Join us to improve resilience and delivery speed while keeping downtime low

Responsibilities
  • Design, build and maintain cloud infrastructure using modern IaC practices such as Terraform or CloudFormation
  • Create and optimize CI/CD pipelines to automate software deployments, configuration management and repetitive operational tasks
  • Implement robust logging, monitoring and alerting systems to establish clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
  • Respond to production incidents and lead troubleshooting efforts to restore services
  • Run blameless post-mortems to identify root causes and prevent recurrence
  • Collaborate with software developers to optimize system performance and plan capacity
  • Ensure services can scale to handle growth and traffic spikes
Requirements
  • Proven experience of 5+ years in systems administration, DevOps, or systems-oriented software development
  • Hands-on proficiency in at least one scripting or programming language such as Python, Bash, Go or Rust
  • Solid experience with public cloud providers such as AWS, Azure or GCP and containerization tools such as Docker and Kubernetes
  • Deep understanding of Linux/Unix administration and networking fundamentals such as TCP/IP, DNS and HTTP/SSL/TLS
  • Working knowledge of monitoring and observability tools such as Prometheus, Grafana or Datadog
  • Clear passion for automation, eliminating toil, and building resilient systems that fail gracefully
  • English proficiency at B2 (Upper-Intermediate) level or higher
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

PeoplePlusTech Inc. • Metro Manila

Hybrid
PHP 900,000 - 1,500,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

LTM • Mexico

On-site
PHP 900,000 - 1,300,000
Staff SRE Engineer
Staff SRE Engineer

Stellar Cyber • España

On-site
PHP 5,528,000 - 7,372,000
Lead SRE: Build Resilient Cloud & Automations
Lead SRE: Build Resilient Cloud & Automations

EPAM Systems • Mexico

On-site
PHP 5,846,000 - 8,616,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

AIPI Acquire Intelligence Philippines Inc. • Taguig

On-site
PHP 1,000,000 - 1,500,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Acquire Intelligence • Taguig

On-site
PHP 900,000 - 1,500,000
Technical Lead - Site Reliability Engineering
Technical Lead - Site Reliability Engineering

LSEG • Taguig

On-site
PHP 4,914,000 - 7,372,000
Healthcare
Retirement planning
Paid volunteering days
+1
Associate Site Reliability Engineer
Associate Site Reliability Engineer

Railway Corp • Mexico

On-site
PHP 5,474,000 - 7,908,000
SRE (Site Reliability Engineer)
SRE (Site Reliability Engineer)

GCash • Manila

On-site
Opportunity for career growth and development
Dynamic collaborative team environment
Highly competitive compensation and benefits package
Site Reliability Engineer
Site Reliability Engineer

Alsons/AWS Information Systems Inc. • Cebu City

Hybrid
PHP 600,000 - 1,000,000