Senior Site Reliability Engineer (SRE)

EPAM Systems

Morelia

Presencial

MXN 600.000 - 900.000

Jornada completa

hace 23 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Transforma esta oferta en una entrevista: un currículum y una carta de presentación creados pensando en lo que quiere el empleador.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Healthcare benefits
Paid time off
Upskilling and certification courses
LinkedIn Learning access
Volunteer opportunities
Global career opportunities

Descripción de la vacante

EPAM Systems is seeking a Senior Site Reliability Engineer to bridge software development and operations, automating processes, scaling infrastructure, and ensuring systems are highly available and performant. Your mission is to build, run, and protect production environments that power our applications, minimizing downtime and enabling rapid software deployment.

Responsibilities include designing cloud infrastructure with Terraform/CloudFormation, building CI/CD pipelines, implementing logging

Formación

  • 3+ years in systems administration, DevOps, or related roles.
  • Experience with cloud providers (AWS/Azure/GCP) and containerization (Docker/Kubernetes).
  • Strong automation mindset and ability to build resilient systems.

Responsabilidades

  • Design, build, and maintain cloud infrastructure using IaC (Terraform, CloudFormation).
  • Develop CI/CD pipelines to automate deployments and reduce toil.
  • Implement monitoring/alerting with Prometheus, Grafana, and Datadog.
  • Define SLOs/SLIs and participate in blameless post-mortems.
  • Collaborate with software teams to optimize performance and scalability.
  • Respond to incidents and lead root-cause analysis.

Conocimientos

Strong problem solving
English proficiency

Herramientas

AWS
Azure
GCP
Docker
Kubernetes
Terraform
CloudFormation

Descripción del empleo

EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.

We are seeking a skilled and proactive Senior Site Reliability Engineer (SRE) to join our engineering team. In this role, you will bridge the gap between software development and systems operations, applying software engineering principles to automate operations, scale infrastructure, and ensure systems are highly available, resilient, and performant. Your mission is to build, run, and protect the production environments that power our applications, minimizing downtime and enabling rapid, safe software deployment.

Responsibilities
  • Design, build, and maintain cloud infrastructure using modern Infrastructure as Code practices such as Terraform and CloudFormation
  • Build and optimize CI/CD pipelines to automate software deployments, configuration management, and repetitive operational tasks
  • Design and implement robust logging, monitoring, and alerting systems using tools such as Prometheus, Grafana, and Datadog
  • Establish clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
  • Respond to production incidents and lead troubleshooting efforts to restore services
  • Conduct blameless post-mortems to identify root causes and prevent recurrence
  • Partner with software developers to optimize system performance and plan capacity
  • Ensure services can scale to handle growth and traffic spikes
Requirements
  • 3+ years of experience in systems administration, DevOps, or systems-focused software development
  • Proficiency in at least one scripting or programming language such as Python, Bash, Go, or Rust
  • Experience with public cloud providers such as AWS, Azure, or GCP, along with containerization tools such as Docker and Kubernetes
  • Understanding of Linux/Unix administration and networking fundamentals such as TCP/IP, DNS, and HTTP/SSL/TLS
  • Passion for automation, eliminating toil, and building resilient systems that fail gracefully
  • Advanced proficiency in English (C1+)
We offer
  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn

EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior DevOps Engineer (AWS)
Senior DevOps Engineer (AWS)

EPAM Systems • Morelia

Presencial
MXN 600.000 - 1.000.000
Healthcare benefits
Employee financial programs
Paid time off and sick leave
+4
Senior DevOps Engineer (AWS)
Senior DevOps Engineer (AWS)

EPAM Systems • Región Centro

Presencial
MXN 700.000 - 1.000.000
Healthcare benefits
Professional development
Paid time off
+2
Senior SRE: Build Resilient Infra & Automate Deployments
Senior SRE: Build Resilient Infra & Automate Deployments

EPAM Systems • Morelia

Presencial
MXN 600.000 - 900.000
Healthcare benefits
Paid time off
Upskilling and certification courses
+3
Site Reliability Engineer ID60188
Site Reliability Engineer ID60188

AgileEngine • Ciudad de México

Híbrido
MXN 1.049.000 - 1.400.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Rosarito

Híbrido
MXN 870.000 - 1.306.000
Professional growth
Competitive compensation
Exciting projects
+1
Project Manager
Project Manager

EPAM Systems • Aguascalientes

Presencial
MXN 600.000 - 900.000
Healthcare benefits
Paid time off
LinkedIn Learning access
+3
Site Reliability Engineer
Site Reliability Engineer

Pyramid Consulting, Inc • Estado de México

Presencial
Senior Business Analyst
Senior Business Analyst

EPAM Systems • León

Presencial
MXN 420.000 - 660.000
Chief AI Solution Engineer
Chief AI Solution Engineer

EPAM Systems • Región Centro

Presencial
MXN 1.200.000 - 1.800.000
Healthcare benefits
Paid time off & sick leave
Global career opportunities
+2
Site Reliability Engineer
Site Reliability Engineer

CTC • Estado de México

A distancia
MXN 1.433.000 - 1.793.000