Lead Site Reliability Engineer

EPAM Systems

Mexico

On-site

PHP 5,474,000 - 7,908,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Healthcare benefits
Paid time off
Upskilling courses
LinkedIn Learning
Global career paths
Volunteer opportunities
Employee groups
Award-winning culture

Job summary

EPAM Systems in the Philippines is seeking a Lead Compute Platform SRE to support its Compute Managed Services projects. You will drive 24x7 monitoring, incident management, and operational stability across multi-cloud environments.

Collaborating with cross-functional teams, you will advance observability, automate processes, enforce security and compliance, and help deliver high-quality compute services. This role requires leadership experience, strong English communication, and hands-on

Qualifications

  • Minimum 5 years of relevant experience.
  • At least 1 year of experience leading and managing teams.
  • Experience with cloud platforms such as GCP, AWS, and Azure.
  • Proficient in OS administration across Windows and Linux.
  • Proficient in automation tools such as Ansible, Terraform, Python, and Bash.
  • Strong knowledge of observability tools such as ELK and Grafana.
  • Solid understanding of incident management processes.
  • Experience using GitHub for version control and collaborative development.
  • Experience in security hardening, vulnerability management, and compliance practices.
  • Excellent problem-solving, communication, and collaboration skills.
  • Familiarity with disaster recovery and operational recovery processes.
  • English level B2 or higher, with strong written and verbal communication skills.

Responsibilities

  • Perform continuous 24x7 monitoring of compute platforms using tools such as ELK and PagerDuty
  • Manage incidents and problems across servers, middleware, operating systems, and cloud platforms, including troubleshooting, root cause analysis (RCA), and resolution
  • Execute repaving activities, change management processes, and disaster recovery procedures
  • Ensure security and vulnerability compliance, including user management and certificate lifecycle management
  • Handle service requests, configuration updates, and audit-related data extracts
  • Develop and maintain Standard Operating Procedures (SOPs) for infrastructure operations
  • Collaborate on cell-based automation improvements and drive continuous service enhancements

Skills

Leadership
Cloud platforms (GCP/AWS/Azure)
OS administration (Windows/Linux)
Automation (Ansible, Terraform, Python
Observability (ELK, Grafana)
Incident management
GitHub
Security hardening & compliance
Communication & collaboration

Tools

Ansible
Terraform
Python
Bash
ELK
Grafana
GitHub

Job description

EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.

We are seeking a Lead Compute Platform SRE to support EPAM's Compute Managed Services project for our client.

The role focuses on KTLO (Keep the Lights On) activities, ensuring 24x7 monitoring, incident management, and operational stability across multi-cloud environments (GCP, AWS, Azure). The SRE will drive observability improvements, automate processes, and maintain compliance while collaborating with cross-functional teams to deliver high-quality compute services.

Responsibilities
  • Perform continuous 24x7 monitoring of compute platforms using tools such as ELK and PagerDuty
  • Manage incidents and problems across servers, middleware, operating systems, and cloud platforms, including troubleshooting, root cause analysis (RCA), and resolution
  • Execute repaving activities, change management processes, and disaster recovery procedures
  • Ensure security and vulnerability compliance, including user management and certificate lifecycle management
  • Handle service requests, configuration updates, and audit-related data extracts
  • Develop and maintain Standard Operating Procedures (SOPs) for infrastructure operations
  • Collaborate on cell-based automation improvements and drive continuous service enhancements
Requirements
  • A minimum of 5 years of relevant experience
  • At least one year of experience leading and managing teams
  • Experience working with cloud platforms such as GCP, AWS, and Azure
  • Proficiency in OS administration across Windows and Linux environments
  • Proficiency in automation tools such as Ansible, Terraform, Python, and Bash
  • Strong knowledge of observability tools such as ELK and Grafana
  • Solid understanding of incident management processes
  • Experience using GitHub for version control and collaborative development
  • Experience in security hardening, vulnerability management, and compliance practices
  • Excellent problem-solving, communication, and collaboration skills
  • Familiarity with disaster recovery and operational recovery processes
  • English level B2 or higher, with strong written and verbal communication skills
We offer
  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn

EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Delivery Manager
Data Delivery Manager

EPAM Systems • Mexico

On-site
PHP 6,695,000 - 9,738,000
Healthcare benefits
Employee financial programs
Unlimited LinkedIn Learning access
Senior Data Software Engineer (Python & SQL)
Senior Data Software Engineer (Python & SQL)

EPAM Systems • Mexico

Hybrid
PHP 3,685,000 - 4,915,000
International projects with top brands
Paid time off and sick leave
Unlimited access to the LinkedIn Learning library
+1
Senior Quality Engineer (JavaScript)
Senior Quality Engineer (JavaScript)

EPAM Systems • Mexico

On-site
PHP 5,474,000 - 7,908,000
Healthcare benefits
Global career opportunities
LinkedIn Learning access
Product Owner
Product Owner

EPAM Systems • Mexico

On-site
PHP 5,474,000 - 7,299,000
Healthcare benefits
Employee programs
Paid time off
Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

EPAM Systems • Mexico

On-site
PHP 5,846,000 - 8,616,000
Technical Lead - Site Reliability Engineering
Technical Lead - Site Reliability Engineering

LSEG • Taguig

On-site
PHP 4,914,000 - 7,372,000
Healthcare
Retirement planning
Paid volunteering days
+1
Architecture Technology Consultant
Architecture Technology Consultant

EPAM Systems • Mexico

On-site
PHP 5,230,000 - 6,770,000
Paid time off and sick leave
Access to LinkedIn Learning library
Employee financial programs
Site Reliability Engineer
Site Reliability Engineer

PeoplePlusTech Inc. • Metro Manila

Hybrid
PHP 900,000 - 1,500,000
QA Architect
QA Architect

EPAM Systems • Mexico

On-site
PHP 7,353,000 - 11,029,000
Healthcare benefits
Global career opportunities
Paid time off
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Acquire Intelligence • Taguig

On-site
PHP 900,000 - 1,500,000