Senior Site Reliability Engineer

MPS Group LLC

Bogotá ciudad

Presencial

COP 140.000.000 - 240.000.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Consigue una respuesta de este empleador — un currículum y una carta de presentación adaptados exactamente a lo que busca para contratar.

Supera los filtros ATS

Descripción de la vacante

MPS Group LLC is seeking a Senior Site Reliability Engineer to design, operate, and evolve highly scalable cloud platforms supporting enterprise apps. You will partner with Engineering, DevOps, Platform, and Security teams to embed reliability practices and ensure performance, availability, and resilience.

This hands-on leadership role emphasizes cloud infrastructure, Kubernetes, observability, incident management, and reliability engineering.

Formación

  • 7+ years of SRE/DevOps/Platform Engineering experience.
  • Experience with AWS and/or GCP production environments.
  • Experience with Kubernetes platforms (EKS/GKE).
  • Strong scripting in Python or Bash.
  • Experience with incident management and reliability testing preferred.
  • Familiarity with CI/CD pipelines and IaC.

Responsabilidades

  • Design reliability strategies for distributed systems on AWS and GCP.
  • Define SLIs, SLOs, and reliability metrics.
  • Build observability using monitoring, logging, tracing, and alerting.
  • Lead incident response and postmortems to improve reliability.
  • Collaborate to improve performance, resiliency, and scalability.
  • Automate ops and reduce toil with engineering solutions.
  • Guide architecture decisions for reliability and capacity planning.

Conocimientos

SRE fundamentals
AWS
GCP
Kubernetes
Python
Bash
CI/CD
Observability concepts
Networking basics

Herramientas

Terraform
EKS
GKE
Prometheus
Grafana
CloudWatch
Datadog
Istio

Descripción del empleo

At MPS Group, we empower businesses with cutting-edge technology solutions and expert consulting. We proudly serve a diverse range of industries, including IT, Fintech, Airlines, Energy, and more. Our expertise spans front-end technologies (React, Angular, Vue.js), back-end technologies (Java, .NET, Node.js, Python), and comprehensive Data Solutions (Database, Data Warehousing, BI, Reporting, Analytics). Our services include custom software development, QA Testing (SDET), Automation Testing, IT staffing, business development consulting, and DevOps services. We adopt both Agile and Waterfall methodologies, offering flexibility through onshore, nearshore, and offshore delivery models. From cloud solutions (Microsoft Azure, AWS, GCP) to digital advisory and project management, we craft innovative, tailored strategies designed to meet your specific business goals and drive long-term success.

About the Role

We're looking for a Senior Site Reliability Engineer to help build, operate, and evolve highly scalable, resilient, and secure cloud platforms supporting critical enterprise applications. As part of a large-scale cloud transformation initiative, you will partner closely with Engineering, DevOps, Platform, and Security teams to establish reliability practices, improve operational excellence, and ensure systems meet performance, availability, and scalability objectives. This is a hands-on technical leadership role requiring deep expertise in cloud infrastructure, Kubernetes, observability, incident management, and reliability engineering. You will drive technical decisions, influence engineering practices, and help teams design systems that are resilient by design. Note: This project involves a migration from AWS to GCP, so the candidate must have strong experience with AWS and at least some exposure to GCP.

Responsibilities
  • Design and implement reliability strategies for distributed systems running across AWS and GCP.
  • Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability metrics.
  • Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
  • Lead incident response, root cause analysis, and postmortem processes to improve system reliability.
  • Collaborate with engineering teams to improve system performance, resiliency, scalability, and operational readiness.
  • Automate operational processes and reduce toil through engineering solutions.
  • Guide teams on reliability-focused architecture decisions, capacity planning, and non-functional requirements.
Requirements
  • 7+ years of experience in Site Reliability Engineering, Cloud Engineering, DevOps, or Platform Engineering.
  • Strong experience supporting production systems in AWS and/or GCP environments.
  • Deep understanding of SRE principles, including SLIs, SLOs, error budgets, and operational excellence.
  • Experience operating and troubleshooting Kubernetes platforms such as EKS and/or GKE.
  • Strong knowledge of observability tools such as Prometheus, Grafana, CloudWatch, Cloud Monitoring, Datadog, Splunk, or similar.
  • Experience with Infrastructure as Code tools such as Terraform.
  • Strong scripting and automation skills using Python, Bash, or comparable languages.
  • Solid understanding of networking, distributed systems, cloud security, and performance optimization.
  • Experience supporting large-scale cloud migration or modernization programs (preferred).
  • Expertise in incident management and production operations for high-availability systems (preferred).
  • Experience implementing chaos engineering or resilience testing practices (preferred).
  • Knowledge of service mesh technologies such as Istio (preferred).
  • AWS and/or GCP certifications (preferred).
  • Experience working in Agile, DevOps, or DevSecOps environments (preferred).

Details regarding benefits, compensation packages, and perks will be discussed during the hiring process.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Cloud SRE: AWS/GCP Reliability Leader
Senior Cloud SRE: AWS/GCP Reliability Leader

MPS Group LLC • Bogotá ciudad

Presencial
COP 140.000.000 - 240.000.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LanceSoft, Inc. • Colombia

Presencial
COP 90.000.000 - 150.000.000
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Metropolitana

Híbrido
COP 142.369.000 - 213.554.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Adidas • Bogotá ciudad, Guavio

Híbrido
COP 9.000.000 - 13.000.000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

FashionUnited Group • Bogotá ciudad

Híbrido
COP 144.000.000 - 216.000.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

N-iX • Colombia

Híbrido
COP 120.000.000 - 180.000.000
Flexible work format
Education reimbursement
Professional development
Site Reliability Engineer ID62591
Site Reliability Engineer ID62591

AgileEngine • Metropolitana

Híbrido
COP 149.902.000 - 224.854.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Site Reliability Engineer
Site Reliability Engineer

Intraway • Bogotá

Presencial
COP 96.000.000 - 140.000.000
Unlimited PTO
Training access
English classes
+2
Senior DevOps Engineer ID56470
Senior DevOps Engineer ID56470

AgileEngine • Bogotá

Híbrido
COP 285.938.000 - 428.909.000
Professional growth
Competitive compensation
Exciting projects
+1
Senior Cloud Site Reliability Engineer Manager
Senior Cloud Site Reliability Engineer Manager

ScotiaTech • Bogotá

Presencial
COP 180.000.000 - 240.000.000