SRE Software Engineer

Jobtailor

Bogotá

Presencial

COP 100.000.000 - 180.000.000

Jornada completa

14 días+

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

Jobtailor in Bogotá, Colombia is seeking a Site Reliability Engineer to join our operations team. You will support and maintain production environments, manage Kubernetes clusters, and optimize cloud infrastructure across AWS/Azure/GCP.

The role emphasizes reliability, incident response, RCA, and continuous improvement in collaboration with engineering and product teams. You will implement IaC, automate deployments, develop observability dashboards in Prometheus/Grafana, and contribute to

Formación

  • 4+ years of experience in Site Reliability Engineering, DevOps, Cloud Operations, or Infrastructure Engineering.
  • Strong hands-on experience with Linux administration, troubleshooting, and production support.
  • Experience managing and supporting Kubernetes and containerized workloads (Docker/OpenShift is a plus).
  • Solid knowledge of AWS, Azure, or GCP cloud environments.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, Splunk, ELK, or CloudWatch.
  • Experience with Infrastructure as Code (Terraform preferred) and CI/CD pipelines.
  • Ability to troubleshoot complex production issues, perform Root Cause Analysis (RCA), and drive preventive improvements.
  • Working knowledge of automation and scripting using Bash, Python, or Go.
  • Intermediate to advanced English (B2+).

Responsabilidades

  • Support and maintain business-critical production environments, ensuring high availability and system reliability.
  • Monitor infrastructure and applications, proactively identifying and resolving issues before they impact users.
  • Participate in incident response activities, troubleshooting production outages and coordinating recovery efforts.
  • Perform RCA and contribute to postmortems, corrective actions, and continuous improvement initiatives.
  • Manage and optimize Kubernetes clusters and cloud infrastructure.
  • Develop and maintain monitoring dashboards, alerts, and observability solutions.
  • Automate operational processes and infrastructure deployments using IaC and scripting.
  • Collaborate with engineering and product teams to improve scalability, performance, and operational excellence.
  • Support and enhance CI/CD pipelines to ensure reliable and efficient software delivery.

Conocimientos

Site Reliability Engineering
Kubernetes Management
Cloud Infrastructure (AWS, Azure, GCP)
Infrastructure as Code (Terraform)
Monitoring & Observability (Prometheus
Grafana & Datadog
CI/CD Pipelines
Automation and Scripting (Bash, Python

Herramientas

Kubernetes
Docker
OpenShift
Prometheus
Grafana
Datadog
Splunk
ELK
CloudWatch

Descripción del empleo

  • Support and maintain business-critical production environments, ensuring high availability and system reliability.
  • Monitor infrastructure and applications, proactively identifying and resolving issues before they impact users.
  • Participate in incident response activities, troubleshooting production outages and coordinating recovery efforts.
  • Perform RCA and contribute to postmortems, corrective actions, and continuous improvement initiatives.
  • Manage and optimize Kubernetes clusters and cloud infrastructure.
  • Develop and maintain monitoring dashboards, alerts, and observability solutions.
  • Automate operational processes and infrastructure deployments using IaC and scripting.
  • Collaborate with engineering and product teams to improve scalability, performance, and operational excellence.
  • Support and enhance CI/CD pipelines to ensure reliable and efficient software delivery.
Requirements
  • 4+ years of experience in Site Reliability Engineering, DevOps, Cloud Operations, or Infrastructure Engineering.
  • Strong hands-on experience with Linux administration, troubleshooting, and production support.
  • Experience managing and supporting Kubernetes and containerized workloads (Docker/OpenShift is a plus).
  • Solid knowledge of AWS, Azure, or GCP cloud environments.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, Splunk, ELK, or CloudWatch.
  • Experience with Infrastructure as Code (Terraform preferred) and CI/CD pipelines.
  • Ability to troubleshoot complex production issues, perform Root Cause Analysis (RCA), and drive preventive improvements.
  • Working knowledge of automation and scripting using Bash, Python, or Go.
  • Intermediate to advanced English (B2+).
Core Competencies

Demonstrates expertise in Site Reliability Engineering and DevOps practices, with a strong focus on managing Kubernetes clusters, cloud infrastructure, and automation through Infrastructure as Code. Proficient in monitoring and observability tools to ensure high availability and system reliability.

Highest-signal resume keywords
  • Site Reliability Engineering
  • Kubernetes Management
  • Cloud Infrastructure (AWS, Azure, GCP)
  • Infrastructure as Code (Terraform)
  • Monitoring and Observability Tools (Prometheus, Grafana, Datadog)
ATS Optimization Keywords
Hard Skills
  • Linux Administration
  • Troubleshooting
  • Root Cause Analysis (RCA)
  • Automation and Scripting (Bash, Python, Go)
  • CI/CD Pipelines
Industry Keywords
  • DevOps
  • Cloud Operations
  • Infrastructure Engineering
  • Production Support
  • Continuous Improvement
Tools & Technologies
  • Kubernetes
  • Docker
  • OpenShift
  • Prometheus
  • Grafana
  • Datadog
  • Splunk
  • ELK
  • CloudWatch
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

SRE Software Engineer
SRE Software Engineer

Capgemini Engineering • Bogotá

Presencial
COP 90.000.000 - 150.000.000
Intermediate SRE
Intermediate SRE

CBTW Americas • Colombia

Presencial
COP 66.960.000 - 100.440.000
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Metropolitana

Híbrido
COP 142.369.000 - 213.554.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Bilingual Linux Infrastructure SRE
Bilingual Linux Infrastructure SRE

Capgemini Engineering • Colombia

Presencial
COP 288.646.000 - 481.078.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

LanceSoft, Inc. • Colombia

Presencial
COP 90.000.000 - 150.000.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

MPS Group LLC • Bogotá

Presencial
COP 200.880.000 - 312.480.000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Publicis Sapient • Colombia

Presencial
COP 284.270.000 - 473.784.000
SRE Software Engineer
SRE Software Engineer

Capgemini • Bogotá

Híbrido
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Oowlish • Bogotá

Presencial
COP 180.000.000 - 300.000.000
Home office
Career plans to allow for extensive成长
International projects
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Bogotá

Presencial
COP 244.056.000 - 418.382.000