Site Reliability Engineer

AgileEngine

Colombia

Presencial

COP 283.921.000 - 441.654.000

Jornada completa

Hace 2 días
Sé de los primeros/as/es en solicitar esta vacante

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Ventajas ofrecidas por este puesto de trabajo

Professional growth
Competitive USD compensation
Flextime
Mentorship & TechTalks

Descripción de la vacante

AgileEngine seeks a DevOps / Site Reliability Engineer to maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. You will lead major-incident calls, own remediation, and build playbooks that guide response.

The role blends platform engineering with incident command, using Terraform, CI/CD pipelines, and CSPM tools like Wiz to secure workloads and automate runbooks while mentoring mid-level SREs.

Formación

  • 5+ years of experience in DevOps/SRE.
  • Expert in multi-cloud defense, IAM, and zero-trust principles.
  • Hands-on with Kubernetes, Terraform, CI/CD orchestration, and Python/Go.
  • Senior-level incident-command experience in 24x7 production environments.
  • Proven remediation follow-up and cross-team coordination.
  • Experience with CNAPP/CSPM platforms, Wiz preferred.
  • Familiar with PCI-DSS and SOC2 compliance.

Responsabilidades

  • Scale and maintain operational stability across Azure, AWS, and GCP.
  • Engineer unified security policies using IaC (Terraform).
  • Design and optimize enterprise CI/CD pipelines.
  • Respond to monitoring alerts and secure workloads with CSPM tools.
  • Act as Incident Commander for major incidents and coordinate teams.
  • Own the post-incident closure and systemic fixes.
  • Draft incident notifications to technical and executive audiences.
  • Develop and socialize incident-management playbooks and runbooks.

Conocimientos

Kubernetes
Terraform
CI/CD orchestration
Python/Go scripting
Incident-command experience
Multi-cloud defense
Zero-trust principles
PCI-DSS
SOC2
Wiz CSPM

Herramientas

Wiz
PagerDuty
ServiceNow

Descripción del empleo

We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. This role blends platform engineering with incident command, using Terraform, CI/CD pipelines, and CSPM tools like Wiz. You will lead major-incident calls, own remediation follow-through, and build the playbooks that guide response.

What you will do
  • Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP).
  • Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations.
  • Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment.
  • Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads.
  • Serve as Incident Commander on major and critical incidents — running the bridge, directing technical workstreams, making time-critical decisions, and coordinating cross-functional responders under pressure.
  • Own the post-incident loop — track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups.
  • Draft and send clear, accurate, audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle.
  • Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution.
Must haves
  • 5+ years of experience.
  • In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles.
  • Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting.
  • Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment.
  • Proven track record of remediation follow-up — coordinating with teams and holding owners accountable until issues are fully closed.
  • Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences.
  • Direct experience authoring divisional/group incident-management playbooks and escalation procedures.
  • Fully autonomous.
  • Drives the architecture of complex automated runbooks and mentors Middle-level SREs.
  • Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz.
  • Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2).
Nice to haves
  • PagerDuty — hands-on experience with on-call scheduling, alert routing, and incident orchestration.
  • ServiceNow — familiarity with incident, problem, and change management workflows and reporting.
Perks and Benefits
  • Professional growth

Accelerate your professional journey with mentorship, TechTalks, and personalized growth roadmaps

  • Competitive compensation

We match your ever-growing skills, talent, and contributions with competitive USD-based compensation and budgets for education, fitness, and team activities

  • A selection of exciting projects

Join projects with modern solutions development and top-tier clients that include Fortune 500 enterprises and leading product brands

  • Flextime

Tailor your schedule for an optimal work-life balance, by having the options of working from home and going to the office – whatever makes you the happiest and most productive.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

DevOps / Site Reliability Engineer ID70127
DevOps / Site Reliability Engineer ID70127

AgileEngine, LLC. • Centrosur

Presencial
COP 100.440.000 - 167.400.000
Growth without limits
Competitive compensation
Flexibility: 100% remote with flexible
+3
DevOps / Site Reliability Engineer ID70127
DevOps / Site Reliability Engineer ID70127

AgileEngine, LLC. • Sur

Presencial
COP 120.000.000 - 240.000.000
Growth opportunities
Competitive pay
Remote-friendly schedule
+3
DevOps / Site Reliability Engineer ID70127
DevOps / Site Reliability Engineer ID70127

AgileEngine, LLC. • Metropolitana

Presencial
COP 120.000.000 - 240.000.000
Growth opportunities
Competitive pay
Remote work
+3
Senior Multi-Cloud SRE & Incident Commander
Senior Multi-Cloud SRE & Incident Commander

AgileEngine • Colombia

Presencial
COP 283.921.000 - 441.654.000
Professional growth
Competitive USD compensation
Flextime
+1
Remote Senior DevOps & SRE — Multi-Cloud & Incident Command
Remote Senior DevOps & SRE — Multi-Cloud & Incident Command

AgileEngine, LLC. • Centrosur

Presencial
COP 100.440.000 - 167.400.000
Growth without limits
Competitive compensation
Flexibility: 100% remote with flexible
+3
Remote Senior DevOps & SRE — Multi-Cloud Incident Lead
Remote Senior DevOps & SRE — Multi-Cloud Incident Lead

AgileEngine, LLC. • Perímetro Urbano Barranquilla

Presencial
COP 167.400.000 - 290.160.000
Growth opportunities
Competitive salary
Remote-friendly
+3
Senior DevOps & SRE - Multi-Cloud Incident Lead
Senior DevOps & SRE - Multi-Cloud Incident Lead

AgileEngine, LLC. • Pereira

Presencial
COP 90.000.000 - 170.000.000
Growth opportunities
Competitive pay
Remote work
+3
Senior DevOps & SRE - Multi-Cloud Incident Lead (Remote)
Senior DevOps & SRE - Multi-Cloud Incident Lead (Remote)

AgileEngine, LLC. • Metropolitana

Presencial
COP 120.000.000 - 240.000.000
Growth opportunities
Competitive pay
Remote work
+3
Senior DevOps & SRE: Multi-Cloud Incident Commander Remote
Senior DevOps & SRE: Multi-Cloud Incident Commander Remote

AgileEngine, LLC. • Cartagena de Indias

Presencial
COP 120.000.000 - 190.000.000
Growth without limits
Competitive compensation
Flexibility: 100% remote
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Metropolitana

Híbrido
COP 142.369.000 - 213.554.000
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1