Site Reliability Engineer

EPAM Systems, Inc.

México

A distancia

MXN 450.000 - 750.000

Jornada completa

Hace 3 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

No envíes un currículum genérico — crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Healthcare benefits
Paid time off
Upskilling courses
Global career opportunities

Descripción de la vacante

EPAM Systems, Inc. is seeking a Site Reliability Engineer to strengthen reliability, observability, and platform operations across cloud and Kubernetes environments in Mexico.

You will improve service health through automation, infrastructure as code, and CI/CD practices to keep critical systems stable and scalable. You will manage production Kubernetes clusters, build CI/CD pipelines with Azure DevOps, and implement observability using MELT signals.

Formación

  • 2+ years in site reliability or DevOps.
  • Kubernetes administration for production workloads.
  • Azure DevOps/Azure Pipelines for CI/CD.
  • IaC with Terraform and Ansible.
  • Scripting for automation tasks.
  • Troubleshooting across MELT: metrics, events, logs, traces.
  • English proficiency at B1+.

Responsabilidades

  • Maintain service reliability by handling L2 operations and incident response.
  • Operate and administer Kubernetes clusters to ensure stability and performance.
  • Build and improve CI/CD workflows using Azure DevOps and Azure Pipelines.
  • Automate operational tasks using scripting to reduce manual effort.
  • Define and maintain infrastructure as code using Terraform and Ansible.
  • Implement observability using MELT signals to detect and resolve issues faster.
  • Coordinate problem resolution by analyzing root causes and proposing corrective actions.
  • Support secure and resilient cloud operations on Microsoft Azure.
  • Document operational procedures and share knowledge to improve support readiness.

Conocimientos

Kubernetes administration
Azure DevOps
Terraform
Ansible
Scripting
Observability (MELT)
Incident response
English proficiency

Herramientas

Grafana
Elastic Stack
Apache Cassandra
HashiCorp Vault

Descripción del empleo

We are looking for a Site Reliability Engineer to strengthen reliability, observability, and platform operations across cloud and Kubernetes environments. In this role, you will improve service health through automation, infrastructure as code and CI/CD practices. Apply now to help keep critical systems stable and scalable!ResponsibilitiesMaintain service reliability by handling L2 operations and incident responseOperate and administer Kubernetes clusters to ensure stability and performanceBuild and improve CI/CD workflows using Azure DevOps and Azure PipelinesAutomate operational tasks using scripting to reduce manual effortDefine and maintain infrastructure as code using Terraform and AnsibleImplement and refine observability using MELT signals to detect and resolve issues fasterCoordinate problem resolution by analyzing root causes and proposing corrective actionsSupport secure and resilient cloud operations on Microsoft AzureDocument operational procedures and share knowledge to improve support readinessRequirements2+ years of site reliability engineering or DevOps experienceKubernetes administration experience supporting production workloadsAzure DevOps and Azure Pipelines experience delivering CI/CD workflowsInfrastructure as Code expertise with Terraform and AnsibleProficiency in scripting languages for automation tasksStrong troubleshooting skills across metrics, events, logs, and traces (MELT)Strong understanding of observability concepts and toolsAzure fundamentals knowledge with AZ-900 or AZ-104 certification (or higher)Good communication skills for cross-team incident coordinationEnglish proficiency: B1+ level or higherNice to haveArgo CD administration or implementation experienceExperience with Apache Cassandra cluster or timeseries/NoSQL operationsFamiliarity with Grafana and Elastic Cloud or Elastic StackHashiCorp Vault experienceKnowledge of programming languages, such as Python, Angular, or GoWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedInEPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Site Reliability Engineer
Senior Site Reliability Engineer

EPAM Systems, Inc. • México

A distancia
MXN 350.000 - 550.000
Healthcare benefits
Paid time off and sick leave
Upskilling and cert courses
+1
Site Reliability Engineer: Cloud Kubernetes & CI/CD Expert
Site Reliability Engineer: Cloud Kubernetes & CI/CD Expert

EPAM Systems, Inc. • México

A distancia
MXN 450.000 - 750.000
Healthcare benefits
Paid time off
Upskilling courses
+1
Lead Site Reliability Engineer
Lead Site Reliability Engineer

EPAM Systems, Inc. • México

A distancia
MXN 900.000 - 1.300.000
Lead Data DevOps
Lead Data DevOps

EPAM Systems, Inc. • México

A distancia
MXN 900.000 - 1.300.000
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Rosarito

Presencial
MXN 870.019 - 1.305.028
Professional growth
Competitive compensation
Exciting projects
+1
Senior AWS DevOps Engineer
Senior AWS DevOps Engineer

EPAM Systems, Inc. • México

A distancia
MXN 900.000 - 1.350.000
Healthcare benefits
Paid time off and sick leave
Upskilling, reskilling and certificate
+4
Senior Data DevOps
Senior Data DevOps

EPAM Systems, Inc. • México

A distancia
MXN 900.000 - 1.500.000
Healthcare benefits
Paid time off
LinkedIn Learning access
+2
Site Reliability Engineer ID60188
Site Reliability Engineer ID60188

AgileEngine • Ciudad de México

Presencial
MXN 1.049.685 - 1.399.580
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Lead AWS DevOps Engineer
Lead AWS DevOps Engineer

EPAM Systems, Inc. • México

A distancia
MXN 850.000 - 1.100.000
Healthcare benefits
Paid time off
Upskilling and certification courses
+2
Senior Infrastructure Engineer
Senior Infrastructure Engineer

EPAM Systems, Inc. • México

A distancia
MXN 900.000 - 1.200.000
Healthcare benefits
Paid time off
Career development