SRE Software Engineer

Capgemini Engineering

Bogotá

Presencial

COP 90.000.000 - 150.000.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

Capgemini Engineering in Bogotá seeks an experienced Site Reliability Engineer/DevOps professional to join our production support team. You will monitor, maintain and optimize critical production environments, implement IaC and CI/CD, and collaborate with cloud and engineering colleagues to improve reliability and performance.

Requirements include 4+ years in SRE/DevOps, strong Linux, Kubernetes, cloud knowledge (AWS/Azure/GCP), plus English at B2+.

Formación

  • 4+ years of experience in Site Reliability Engineering, DevOps, Cloud Operations, or Infrastructure Engineering.
  • Strong hands-on experience with Linux administration, troubleshooting, and production support.
  • Experience managing and supporting Kubernetes and containerized workloads (Docker/OpenShift is a plus).
  • Solid knowledge of AWS, Azure, or GCP cloud environments.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, Splunk, ELK, or CloudWatch.
  • Experience with Infrastructure as Code (Terraform preferred) and CI/CD pipelines.
  • Ability to troubleshoot complex production issues, perform Root Cause Analysis (RCA), and drive preventive improvements.
  • Working knowledge of automation and scripting using Bash, Python, or Go.
  • Intermediate to advanced English (B2+).

Responsabilidades

  • Support and maintain business-critical production environments, ensuring high availability and system reliability.
  • Monitor infrastructure and applications, proactively identifying and resolving issues before they impact users.
  • Participate in incident response activities, troubleshooting production outages and coordinating recovery efforts.
  • Perform RCA and contribute to postmortems, corrective actions, and continuous improvement initiatives.
  • Manage and optimize Kubernetes clusters and cloud infrastructure.
  • Develop and maintain monitoring dashboards, alerts, and observability solutions.
  • Automate operational processes and infrastructure deployments using IaC and scripting.
  • Collaborate with engineering and product teams to improve scalability, performance, and operational excellence.
  • Support and enhance CI/CD pipelines to ensure reliable and efficient software delivery

Conocimientos

Site Reliability Engineering
DevOps
Cloud Operations
Infrastructure Engineering
Linux administration
Root Cause Analysis (RCA)
Automation scripting
Go / Python / Bash
English (B2+)

Herramientas

Docker
Kubernetes
OpenShift
Terraform
CI/CD pipelines
Prometheus
Grafana
Datadog
Splunk
ELK
CloudWatch

Descripción del empleo

  • 4+ years of experience in Site Reliability Engineering, DevOps, Cloud Operations, or Infrastructure Engineering.
  • Strong hands‑on experience with Linux administration, troubleshooting, and production support.
  • Experience managing and supporting Kubernetes and containerized workloads (Docker/OpenShift is a plus).
  • Solid knowledge of AWS, Azure, or GCP cloud environments.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, Splunk, ELK, or CloudWatch.
  • Experience with Infrastructure as Code (Terraform preferred) and CI/CD pipelines.
  • Ability to troubleshoot complex production issues, perform Root Cause Analysis (RCA), and drive preventive improvements.
  • Working knowledge of automation and scripting using Bash, Python, or Go.
  • Intermediate to advanced English (B2+).
Job Description
Your Profile
  • 4+ years of experience in Site Reliability Engineering, DevOps, Cloud Operations, or Infrastructure Engineering.
  • Strong hands‑on experience with Linux administration, troubleshooting, and production support.
  • Experience managing and supporting Kubernetes and containerized workloads (Docker/OpenShift is a plus).
  • Solid knowledge of AWS, Azure, or GCP cloud environments.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, Splunk, ELK, or CloudWatch.
  • Experience with Infrastructure as Code (Terraform preferred) and CI/CD pipelines.
  • Ability to troubleshoot complex production issues, perform Root Cause Analysis (RCA), and drive preventive improvements.
  • Working knowledge of automation and scripting using Bash, Python, or Go.
  • Intermediate to advanced English (B2+).
Responsibilities
  • Support and maintain business-critical production environments, ensuring high availability and system reliability.
  • Monitor infrastructure and applications, proactively identifying and resolving issues before they impact users.
  • Participate in incident response activities, troubleshooting production outages and coordinating recovery efforts.
  • Perform RCA and contribute to postmortems, corrective actions, and continuous improvement initiatives.
  • Manage and optimize Kubernetes clusters and cloud infrastructure.
  • Develop and maintain monitoring dashboards, alerts, and observability solutions.
  • Automate operational processes and infrastructure deployments using IaC and scripting.
  • Collaborate with engineering and product teams to improve scalability, performance, and operational excellence.
  • Support and enhance CI/CD pipelines to ensure reliable and efficient software delivery
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

SRE Software Engineer
SRE Software Engineer

Jobtailor • Bogotá

Presencial
COP 100.000.000 - 180.000.000
SRE Software Engineer
SRE Software Engineer

Capgemini • Bogotá

Híbrido
Bilingual Linux Infrastructure SRE
Bilingual Linux Infrastructure SRE

Capgemini Engineering • Colombia

Presencial
COP 288.646.000 - 481.078.000
Intermediate SRE
Intermediate SRE

CBTW Americas • Colombia

Presencial
COP 66.960.000 - 100.440.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LanceSoft, Inc. • Colombia

Presencial
COP 90.000.000 - 150.000.000
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Metropolitana

Híbrido
COP 142.369.000 - 213.554.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MPS Group LLC • Bogotá

Presencial
COP 200.880.000 - 312.480.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Publicis Sapient • Colombia

Presencial
COP 284.270.000 - 473.784.000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Oowlish • Bogotá

Presencial
COP 180.000.000 - 300.000.000
Home office
Career plans to allow for extensive成长
International projects
+3
Senior DevOps Engineer ID56470
Senior DevOps Engineer ID56470

AgileEngine • Bogotá

Híbrido
COP 285.938.000 - 428.909.000
Professional growth
Competitive compensation
Exciting projects
+1