Senior Site Reliability Engineer

Publicis Sapient

Colombia

Presencial

COP 284.270.000 - 473.784.000

Jornada completa

Hace 13 días

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

Publicis Sapient is seeking a Senior Site Reliability Engineer to design and operate scalable cloud platforms for enterprise apps. You will partner with Engineering, DevOps, Platform, and Security to establish reliability practices, improve operational excellence, and meet performance and availability goals.

This hands-on role requires deep expertise in cloud infrastructure, Kubernetes, observability, incident management, and reliability engineering.

Formación

  • 7+ years in Site Reliability Engineering, Cloud or Platform Engineering.
  • Production experience in AWS and/or GCP.
  • Deep knowledge of SRE principles: SLIs, SLOs, error budgets, reliability.
  • Experience operating Kubernetes platforms (EKS/GKE).
  • Strong observability skills with Prometheus, Grafana, CloudWatch, or similar.
  • Terraform or IaC for infrastructure provisioning.
  • Scripting in Python or Bash for automation.

Responsabilidades

  • Design reliability strategies for distributed systems across AWS and GCP.
  • Define, measure and monitor SLIs, SLOs, and reliability metrics.
  • Build and enhance observability with monitoring, logging, tracing, alerts.
  • Lead incident response, root cause analysis, and postmortems.
  • Collaborate with engineering to improve performance, resiliency and scalability.
  • Automate operational tasks to reduce toil and manual work.
  • Guide teams on reliability-focused architecture and capacity planning.

Conocimientos

SRE expertise
Cloud engineering
DevOps
Kubernetes
Observability
Infrastructure as Code
Scripting (Python/Bash)

Herramientas

Terraform
Prometheus
Grafana
CloudWatch
Datadog
Splunk
EKS/GKE

Descripción del empleo

Publicis Sapient is seeking a Senior Site Reliability Engineer to help build, operate, and evolve highly scalable, resilient, and secure cloud platforms supporting critical enterprise applications. As part of a large-scale cloud transformation initiative, you will partner closely with Engineering, DevOps, Platform, and Security teams to establish reliability practices, improve operational excellence, and ensure systems meet performance, availability, and scalability objectives.

This is a hands-on technical role requiring deep expertise in cloud infrastructure, Kubernetes, observability, incident management, and reliability engineering. You will drive technical decisions, influence engineering practices, and help teams design systems that are resilient by design.

Your Impact
  • Design and implement reliability strategies for distributed systems running across AWS and GCP.
  • Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability metrics.
  • Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
  • Lead incident response, root cause analysis, and postmortem processes to improve system reliability.
  • Collaborate with engineering teams to improve system performance, resiliency, scalability, and operational readiness.
  • Automate operational processes and reduce toil through engineering solutions.
  • Guide teams on reliability-focused architecture decisions, capacity planning, and non-functional requirements.
Qualifications
Skills & Experience
  • 7+ years of experience in Site Reliability Engineering, Cloud Engineering, DevOps, or Platform Engineering.
  • Strong experience supporting production systems in AWS and/or GCP environments.
  • Deep understanding of SRE principles, including SLIs, SLOs, error budgets, and operational excellence.
  • Experience operating and troubleshooting Kubernetes platforms such as EKS and/or GKE.
  • Strong knowledge of observability tools such as Prometheus, Grafana, CloudWatch, Cloud Monitoring, Datadog, Splunk, or similar.
  • Experience with Infrastructure as Code tools such as Terraform.
  • Strong scripting and automation skills using Python, Bash, or comparable languages.
  • Solid understanding of networking, distributed systems, cloud security, and performance optimization.
Set Yourself Apart With
  • Experience supporting large-scale cloud migration or modernization programs.
  • Expertise in incident management and production operations for high-availability systems.
  • Experience implementing chaos engineering or resilience testing practices.
  • Knowledge of service mesh technologies such as Istio.
  • AWS and/or GCP certifications.
  • Experience working in Agile, DevOps, or DevSecOps environments.
This position is available just for candidates based in LATAM
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Publicis Sapient • Bogotá

Presencial
COP 180.000.000 - 300.000.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MPS Group LLC • Bogotá

Presencial
COP 200.880.000 - 312.480.000
Senior Cloud Reliability Engineer
Senior Cloud Reliability Engineer

Publicis Sapient • Bogotá

Presencial
COP 180.000.000 - 300.000.000
Senior Cloud SRE — AWS/GCP, Observability & Reliability
Senior Cloud SRE — AWS/GCP, Observability & Reliability

Publicis Sapient • Colombia

Presencial
COP 284.270.000 - 473.784.000
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Metropolitana

Híbrido
COP 142.369.000 - 213.554.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LanceSoft, Inc. • Colombia

Presencial
COP 90.000.000 - 150.000.000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Oowlish • Bogotá

Presencial
COP 180.000.000 - 300.000.000
Home office
Career plans to allow for extensive成长
International projects
+3
Senior SRE: Cloud Reliability Lead (AWS/GCP, Kubernetes)
Senior SRE: Cloud Reliability Lead (AWS/GCP, Kubernetes)

MPS Group LLC • Bogotá

Presencial
COP 200.880.000 - 312.480.000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

EPAM Systems • Colombia

Presencial
COP 60.000.000 - 120.000.000
Healthcare benefits
Global career opportunities
Upskilling and certification courses
+1
Site Reliability Engineer ID62591
Site Reliability Engineer ID62591

AgileEngine • Metropolitana

Híbrido
COP 149.902.000 - 224.854.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1