Lead Cloud SRE: Reliability, Automation & Scale

EPAM Systems, Inc.

Morelia

Presencial

MXN 700.000 - 1.100.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Healthcare benefits
Paid time off and sick leave
Upskilling and certification courses
LinkedIn Learning access
Global career opportunities
Volunteer and community involvement

Descripción de la vacante

EPAM Systems, Inc. seeks a Lead Site Reliability Engineer to keep cloud platforms reliable, scalable, and safe through software-driven operations and automation. You will reduce toil, strengthen availability, and protect production for teams delivering solutions across Financial Services, Insurance, and Retail.

You will design, build and maintain cloud infrastructure, implement CI/CD pipelines, and lead incident response with blameless post-mortems to improve resilience and delivery speed.

Formación

  • 5+ years of experience in systems administration, DevOps, or systems-oriented software development.
  • Hands-on scripting or programming in at least one language (Python, Bash, Go or Rust).
  • Strong cloud experience across AWS, Azure or GCP and containerization with Docker & Kubernetes.
  • Solid Linux/Unix administration and networking fundamentals (TCP/IP, DNS, HTTP/SSL/TLS).
  • Experience with monitoring/observability tools (Prometheus, Grafana, Datadog).
  • English proficiency at B2 (Upper-Intermediate) level or higher.

Responsabilidades

  • Design, build and maintain cloud infrastructure using IaC (Terraform, CloudFormation).
  • Create and optimize CI/CD pipelines to automate deployments and config management.
  • Implement robust logging, monitoring and alerting with clear SLOs/SLIs.
  • Respond to production incidents and lead post-mortems to identify root causes.
  • Collaborate with developers to optimize system performance and plan capacity.
  • Ensure services can scale for growth and traffic spikes.

Descripción del empleo

EPAM Systems, Inc. seeks a Lead Site Reliability Engineer to keep cloud platforms reliable, scalable, and safe through software-driven operations and automation. You will reduce toil, strengthen availability, and protect production for teams delivering solutions across Financial Services, Insurance, and Retail.

You will design, build and maintain cloud infrastructure, implement CI/CD pipelines, and lead incident response with blameless post-mortems to improve resilience and delivery speed.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Lead Site Reliability Engineer (SRE)
Lead Site Reliability Engineer (SRE)

EPAM Systems, Inc. • Morelia

Presencial
MXN 700.000 - 1.100.000
Healthcare benefits
Paid time off and sick leave
Upskilling and certification courses
+3
Senior SRE: Cloud Infra, CI/CD & Resilient Systems
Senior SRE: Cloud Infra, CI/CD & Resilient Systems

EPAM Systems • México

Presencial
MXN 900.000 - 1.300.000
International projects
Global teams
Employee financial programs
+5
Senior DevOps & SRE: Multi-Cloud Resilience | Remote
Senior DevOps & SRE: Multi-Cloud Resilience | Remote

AgileEngine, LLC. • Región Centro

Presencial
MXN 600.000 - 1.000.000
Growth opportunities
Competitive pay
Remote work
Senior Cloud SRE & Automation Engineer
Senior Cloud SRE & Automation Engineer

NTT DATA, Inc. • Región Centro

Presencial
MXN 900.000 - 1.300.000
Lead Site Reliability Engineer - Cloud & AI Ops
Lead Site Reliability Engineer - Cloud & AI Ops

JobCubby • Ciudad de México

Presencial
MXN 900.000 - 1.300.000
Flexible work options
Wellness program
Discounts & scholarships
+2
Remote Multi-Cloud SRE | Incident Commander & IaC
Remote Multi-Cloud SRE | Incident Commander & IaC

AgileEngine, LLC. • Rosarito

Presencial
MXN 700.000 - 1.300.000
Growth opportunities
Competitive compensation
Remote-friendly setup
+1
Senior Site Reliability Engineer - Scale, Automate, Resilient Ops
Senior Site Reliability Engineer - Scale, Automate, Resilient Ops

Mastercard • Ciudad de México

Presencial
MXN 1.200.000 - 2.000.000
Remote Site Reliability Engineer — Cloud & Automation Leader
Remote Site Reliability Engineer — Cloud & Automation Leader

AgileEngine • Rosarito

A distancia
MXN 1.531.000 - 2.212.000
Mentorship and TechTalks
Competitive USD-based compensation
Work on modern solutions
+1
Lead Database Architect & Ops | Flexible Schedule
Lead Database Architect & Ops | Flexible Schedule

EPAM Systems • México

Presencial
MXN 900.000 - 1.300.000
Senior Principal SRE Architect: Observability & Reliability
Senior Principal SRE Architect: Observability & Reliability

iCIMS, Inc. • Ciudad de México

Presencial
MXN 1.800.000 - 3.000.000