Lead Site Reliability Engineer (SRE)

EPAM Systems, Inc.

Morelia

Presencial

MXN 700.000 - 1.100.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Healthcare benefits
Paid time off and sick leave
Upskilling and certification courses
LinkedIn Learning access
Global career opportunities
Volunteer and community involvement

Descripción de la vacante

EPAM Systems, Inc. seeks a Lead Site Reliability Engineer to keep cloud platforms reliable, scalable, and safe through software-driven operations and automation. You will reduce toil, strengthen availability, and protect production for teams delivering solutions across Financial Services, Insurance, and Retail.

You will design, build and maintain cloud infrastructure, implement CI/CD pipelines, and lead incident response with blameless post-mortems to improve resilience and delivery speed.

Formación

  • 5+ years of experience in systems administration, DevOps, or systems-oriented software development.
  • Hands-on scripting or programming in at least one language (Python, Bash, Go or Rust).
  • Strong cloud experience across AWS, Azure or GCP and containerization with Docker & Kubernetes.
  • Solid Linux/Unix administration and networking fundamentals (TCP/IP, DNS, HTTP/SSL/TLS).
  • Experience with monitoring/observability tools (Prometheus, Grafana, Datadog).
  • English proficiency at B2 (Upper-Intermediate) level or higher.

Responsabilidades

  • Design, build and maintain cloud infrastructure using IaC (Terraform, CloudFormation).
  • Create and optimize CI/CD pipelines to automate deployments and config management.
  • Implement robust logging, monitoring and alerting with clear SLOs/SLIs.
  • Respond to production incidents and lead post-mortems to identify root causes.
  • Collaborate with developers to optimize system performance and plan capacity.
  • Ensure services can scale for growth and traffic spikes.

Descripción del empleo

We are seeking a Lead Site Reliability Engineer (SRE) to keep our cloud platforms reliable, scalable, and safe through software-driven operations and automation. You will reduce toil, strengthen availability, and protect production for teams delivering solutions across Financial Services, Insurance, and Retail. Join us to improve resilience and delivery speed while keeping downtime low—apply nowResponsibilitiesDesign, build and maintain cloud infrastructure using modern IaC practices such as Terraform or CloudFormationCreate and optimize CI/CD pipelines to automate software deployments, configuration management and repetitive operational tasksImplement robust logging, monitoring and alerting systems to establish clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs)Respond to production incidents and lead troubleshooting efforts to restore servicesRun blameless post-mortems to identify root causes and prevent recurrenceCollaborate with software developers to optimize system performance and plan capacityEnsure services can scale to handle growth and traffic spikesRequirementsProven experience of 5+ years in systems administration, DevOps, or systems-oriented software developmentHands-on proficiency in at least one scripting or programming language such as Python, Bash, Go or RustSolid experience with public cloud providers such as AWS, Azure or GCP and containerization tools such as Docker and KubernetesDeep understanding of Linux/Unix administration and networking fundamentals such as TCP/IP, DNS and HTTP/SSL/TLSWorking knowledge of monitoring and observability tools such as Prometheus, Grafana or DatadogClear passion for automation, eliminating toil, and building resilient systems that fail gracefullyEnglish proficiency at B2 (Upper-Intermediate) level or higherWe offerInternational projects with top brandsWork with global teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedInEPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Lead Cloud SRE: Reliability, Automation & Scale
Lead Cloud SRE: Reliability, Automation & Scale

EPAM Systems, Inc. • Morelia

Presencial
MXN 700.000 - 1.100.000
Healthcare benefits
Paid time off and sick leave
Upskilling and certification courses
+3
Senior SRE: Cloud Infra, CI/CD & Resilient Systems
Senior SRE: Cloud Infra, CI/CD & Resilient Systems

EPAM Systems • México

Presencial
MXN 900.000 - 1.300.000
International projects
Global teams
Employee financial programs
+5
Site Reliability Engineer ID60188
Site Reliability Engineer ID60188

AgileEngine • Ciudad de México

Híbrido
MXN 1.049.000 - 1.400.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Rosarito

Híbrido
MXN 870.000 - 1.306.000
Professional growth
Competitive compensation
Exciting projects
+1
Site Reilability Engineer
Site Reilability Engineer

Hcltech • Ecatepec de Morelos

Presencial
MXN 520.000 - 760.000
Life insurance
Major Medical Expenses Insurance
Minor Medical Expense Insurance
+4
DevOps / Site Reliability Engineer ID70127
DevOps / Site Reliability Engineer ID70127

AgileEngine, LLC. • Monterrey

Presencial
MXN 670.000 - 1.004.000
Growth without limits
Competitive compensation
Flexibility: 100% remote with flexible
+3
DevOps / Site Reliability Engineer ID70127
DevOps / Site Reliability Engineer ID70127

AgileEngine, LLC. • Santiago de Querétaro

Presencial
MXN 772.000 - 924.000
Growth without limits
Competitive compensation
Remote-friendly culture
+3
Site Reliability Engineer
Site Reliability Engineer

Tata Consultancy Services • Ciudad de México

Presencial
Site Reliability Engineer
Site Reliability Engineer

CTC • Estado de México

A distancia
MXN 1.433.000 - 1.793.000
Site Reliability Engineer
Site Reliability Engineer

Pyramid Consulting, Inc • Estado de México

Presencial