Lead Site Reliability Engineer

EPAM Systems, Inc.

México

A distancia

MXN 900.000 - 1.300.000

Jornada completa

Hace 3 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Destaca en este puesto — crea un currículum adaptado y una carta de presentación en aproximadamente un minuto.

Supera los filtros ATS

Descripción de la vacante

EPAM Systems, Inc. is seeking a Lead Compute Platform SRE to support our Compute Managed Services projects in Mexico. You will own 24x7 monitoring, incident response, and operational stability across multi-cloud environments.

The role emphasizes automation, disaster recovery, security hardening, and SOP development while collaborating with cross-functional teams to deliver high-quality compute services. Candidates should have 5+ years in SRE or IT operations, team leadership experience, and

Formación

  • Minimum 5 years of relevant SRE/IT operations experience.
  • Experience leading and managing teams.
  • Proficient with cloud platforms GCP, AWS and Azure.
  • OS administration across Windows and Linux environments.

Responsabilidades

  • 24x7 monitoring of compute platforms using ELK and PagerDuty.
  • Manage incidents and problems across servers, middleware, OS, and cloud platforms with RCA and resolution.
  • Execute repaving activities, change management, and disaster recovery procedures.
  • Ensure security and vulnerability compliance including user management and certificate lifecycle.
  • Develop and maintain SOPs for infrastructure operations.
  • Collaborate on cell-based automation improvements and drive service enhancements.

Conocimientos

SRE leadership
Cloud platforms
GCP
AWS
Azure
Windows OS
Linux OS
Automation
Ansible
Terraform
Python
Bash
ELK
Grafana
Incident management
GitHub
Security hardening
Vulnerability management
Disaster recovery
English Proficiency

Herramientas

ELK
Grafana
GitHub
Ansible
Terraform

Descripción del empleo

We are seeking a skilled Lead Compute Platform SRE to support EPAM's Compute Managed Services project for our client.The role focuses on KTLO (Keep the Lights On) activities, ensuring 24x7 monitoring, incident management, and operational stability across multi-cloud environments (GCP, AWS, Azure). The SRE will drive observability improvements, automate processes, and maintain compliance while collaborating with cross-functional teams to deliver high-quality compute services.ResponsibilitiesPerform continuous 24x7 monitoring of compute platforms using tools such as ELK and PagerDutyManage incidents and problems across servers, middleware, operating systems, and cloud platforms, including troubleshooting, root cause analysis (RCA), and resolutionExecute repaving activities, change management processes, and disaster recovery proceduresEnsure security and vulnerability compliance, including user management and certificate lifecycle managementHandle service requests, configuration updates, and audit-related data extractsDevelop and maintain Standard Operating Procedures (SOPs) for infrastructure operationsCollaborate on cell-based automation improvements and drive continuous service enhancementsRequirementsA minimum of 5 years of relevant experienceAt least one year of experience leading and managing teamsExperience working with cloud platforms such as GCP, AWS, and AzureProficiency in OS administration across Windows and Linux environmentsProficiency in automation tools such as Ansible, Terraform, Python, and BashStrong knowledge of observability tools such as ELK and GrafanaSolid understanding of incident management processesExperience using GitHub for version control and collaborative developmentExperience in security hardening, vulnerability management, and compliance practicesExcellent problem-solving, communication, and collaboration skillsFamiliarity with disaster recovery and operational recovery processesEnglish level B2 or higher, with strong written and verbal communication skillsWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedInEPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Site Reliability Engineer
Senior Site Reliability Engineer

EPAM Systems, Inc. • México

A distancia
MXN 350.000 - 550.000
Healthcare benefits
Paid time off and sick leave
Upskilling and cert courses
+1
Senior Infrastructure Engineer
Senior Infrastructure Engineer

EPAM Systems, Inc. • México

A distancia
MXN 900.000 - 1.200.000
Healthcare benefits
Paid time off
Career development
Lead SRE – Multi-Cloud Compute Platform & Automation
Lead SRE – Multi-Cloud Compute Platform & Automation

EPAM Systems, Inc. • México

A distancia
MXN 900.000 - 1.300.000
Chief DevOps Engineer
Chief DevOps Engineer

EPAM Systems, Inc. • México

A distancia
MXN 1.200.000 - 2.000.000
Healthcare benefits
Employee financial programs
Paid time off
+3
Chief Cloud Engineer (AWS)
Chief Cloud Engineer (AWS)

EPAM Systems, Inc. • México

A distancia
MXN 1.200.000 - 2.400.000
Healthcare benefits
Global career opportunities
Upskilling & certification courses
+2
Senior Cloud Engineer (AWS)
Senior Cloud Engineer (AWS)

EPAM Systems, Inc. • México

Híbrido
MXN 900.000 - 1.500.000
Lead AWS DevOps Engineer
Lead AWS DevOps Engineer

EPAM Systems, Inc. • México

A distancia
MXN 850.000 - 1.100.000
Healthcare benefits
Paid time off
Upskilling and certification courses
+2
AWS Business Operations Specialist
AWS Business Operations Specialist

EPAM Systems, Inc. • México

A distancia
MXN 900.000 - 1.500.000
Healthcare benefits
Upskilling and certification courses
Global career opportunities
Lead Security Engineer
Lead Security Engineer

EPAM Systems, Inc. • México

A distancia
MXN 900.000 - 1.300.000
Healthcare benefits
Paid time off
Upskilling & certification courses
+3
Technical Program Manager (TPM) – GCP Cloud Infrastructure
Technical Program Manager (TPM) – GCP Cloud Infrastructure

EPAM Systems, Inc. • México

A distancia
MXN 900.000 - 1.100.000