Lead Site Reliability Engineer

Masabi

Colombia

A distancia

COP 299.895.036 - 412.355.675

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Supportive learning culture
Flexible work environment
Focus on personal development

Descripción de la vacante

A leading fintech company is seeking a Lead Site Reliability Engineer to enhance system reliability. This remote role in Colombia involves designing reliable systems, contributing to incident response, and mentoring teams. Candidates should have substantial SRE or DevOps experience, particularly in AWS and infrastructure automation. A supportive and collaborative engineering culture awaits applicants eager to drive improvements in a significant role.

Formación

  • Experience in SRE, platform, or DevOps roles where reliability was critical.
  • Proven experience designing production-grade systems for scale.
  • Hands-on experience with infrastructure automation and CI/CD.

Responsabilidades

  • Lead design discussions for reliability and performance.
  • Implement monitoring solutions for incident signals.
  • Contribute to incident response and root cause analysis.

Conocimientos

SRE experience
AWS knowledge
Terraform
CI/CD experience
Monitoring and observability

Herramientas

Grafana
Prometheus
CloudFormation

Descripción del empleo

Lead Site Reliability Engineer

Introducing Masabi

// At Masabi, we’re driving the fare payment revolution, powering the journeys of millions all over the world. We build fare collection platforms that allow riders to seamlessly buy and present tickets for public transport either on their mobile phones, from a ticket machine, or even by tapping their bank card to travel.

Role

Lead Site Reliability Engineer

We’re looking for a Lead Site Reliability Engineer to join our platform team, someone who’s confident working hands‑on with infrastructure, but also ready to shape how we scale and operate as a global team.

You’ll take ownership of key systems, lead cross‑functional work, and help evolve the way we build for performance, reliability, and security. This role is ideal for those who enjoy solving complex problems, improving systems through automation, and supporting others as they grow. It’s a chance to have both technical depth and meaningful influence, while staying close to the work that matters.

Location

This role is available in a remote model to candidates based in Colombia.

What You’ll Be Doing
Build and automate reliable systems
  • Lead design discussions and make key architectural decisions for reliability, scalability, and performance.
  • Establish SRE standards and best practices (IaC patterns, CI/CD maturity, observability, etc.) across teams.
  • Design and manage infrastructure using Terraform and CloudFormation.
  • Build and evolve CI/CD pipelines that support fast, safe, and frequent deployments.
  • Automate manual tasks to reduce operational load and enable faster delivery.
  • Help expand our infrastructure globally, scaling up new environments with care.
Improve visibility, scale and performance
  • Define and maintain SLIs, SLOs, and alerting strategies aligned with user experience.
  • Implement monitoring solutions that give us clear, early signals during incidents.
  • Lead capacity planning and performance tuning as our systems and teams grow.
  • Identify opportunities to improve architecture for resilience and cost‑effectiveness.
Own reliability and incident response
  • Lead or contribute to incident response, root cause analysis, and post‑incident reviews.
  • Design and maintain disaster recovery and failover strategies.
  • Partner with compliance and security teams to meet frameworks like SOC 2 and PCI.
Support others and share your knowledge
  • Collaborate with engineers, architects, and product teams to embed SRE practices from the start and define long‑term platform reliability strategy.
  • Mentor others in areas like observability, incident readiness, and infrastructure‑as‑code.
  • Document systems and processes clearly to support learning and long‑term success.
  • Partake of the on‑call rotation, shared with the team and paid on top of salary.
About You

// You’re an experienced SRE who combines technical depth with curiosity, care, and a desire to make things better for the platform, the team, and the people using our systems.

  • You’ve worked in SRE, platform, or DevOps roles where reliability was business‑critical (24/7).
  • You have proven experience designing and evolving production‑grade systems for scale and resilience.
  • You’re comfortable designing and operating in AWS, with strong knowledge of cloud architecture, networking and security (VPC design, IAM, least privilege).
  • You have hands‑on experience with Terraform, infrastructure automation, and CI/CD systems.
  • You’ve led or contributed to high‑impact projects involving observability, performance, incident command and/or reliability (distributed tracing, log correlation, metrics maturity, etc).
  • You communicate clearly and drive cross‑functional reliability improvements in distributed, async‑first teams.
  • You enjoy helping others grow and value a kind, collaborative engineering culture.
  • You take pride in doing things the right way, but you’re pragmatic and focused on impact.
Nice To Have
  • Familiarity with PCI DSS v4 or similar compliance standards.
  • Experience with container orchestration.
  • AWS certifications.
Tools & Platforms
  • Monitoring & Observability: Grafana, Prometheus, CloudWatch, Pingdom, Kibana.
  • Infrastructure as Code: Terraform, CloudFormation.
  • Configuration Management & Logging: Puppet, Confluent Cloud.
Why Join Masabi?
  • Driven by Purpose – We believe in journeys made simple. The work isn’t always easy, but the best things never are.
  • Encouraged to Accelerate – Masabi is going places and our people are in the driving seat. Whether you’re taking the direct route or exploring new paths, we support your journey.
  • Advancing with Empathy – We put people first and foster a culture of learning, not blame. No matter your cargo, we share the load.
We’re already powering journeys – are you ready to join us?
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Test Automation Engineer
Test Automation Engineer

Masabi • Colombia

A distancia
COP 132.505.000 - 208.223.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Bogotá

Presencial
COP 244.056.000 - 418.382.000
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Metropolitana

Híbrido
COP 142.369.000 - 213.554.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LanceSoft, Inc. • Colombia

Presencial
COP 90.000.000 - 150.000.000
Site Reliability Engineer
Site Reliability Engineer

DCT • Bogotá

A distancia
COP 156.225.000 - 234.339.000
Career Growth & Mentorship
Flexible Work Environment
Generative & Collaborative Culture
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Publicis Sapient • Colombia

Presencial
COP 284.270.000 - 473.784.000
Platform Operations Engineer (SRE) Colombia / Mexico
Platform Operations Engineer (SRE) Colombia / Mexico

Forte Group • Colombia

Híbrido
COP 226.714.000 - 302.287.000
Site Reliability Engineer ID62591
Site Reliability Engineer ID62591

AgileEngine • Metropolitana

Híbrido
COP 149.902.000 - 224.854.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Service Reliability Engineer
Service Reliability Engineer

1083 Amadeus IT Group Colombia, S.A.S. • Colombia

Presencial
COP 182.089.000 - 254.926.000
Competitive remuneration
Vacation and holiday paid time off
Health insurances
+3
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Oowlish • Bogotá

Presencial
COP 180.000.000 - 300.000.000
Home office
Career plans to allow for extensive成长
International projects
+3