Production Services Lead

Chubb Insurance Hong Kong

Colombia

Presencial

COP 180.000.000 - 300.000.000

Jornada completa

Hace 10 días
Generador de candidaturas

No envíes un currículum genérico — crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Descripción de la vacante

Chubb Insurance Hong Kong seeks a senior SRE leader to run NA Commercial Insurance Production Services and the Site Reliability Engineering function from Bogota. You will ensure reliability, availability and performance of critical production systems, bridging software engineering and operations while mentoring junior engineers.

You will define SLAs/SLOs, drive incidents post-mortems, and collaborate with development, testing and business teams to embed reliability in the SDLC, aiming for 99.9%

Formación

  • 10+ years in SRE, platform engineering, or production operations.
  • 5+ years in a lead or senior individual contributor capacity.
  • Strong reasoning, analytical thinking and troubleshooting skills.
  • Experience with observability and monitoring tools (ELK, App Insights, Splunk, Kibana, AppDynamics, DynaTrace).
  • Basic/intermediate knowledge of MS SQL Server.
  • Strong communication skills and ability to work under pressure.

Responsabilidades

  • Lead a team of SREs and mentor engineers on reliability practices.
  • Own SLAs, SLOs, SLIs and incident response across CI portfolio.
  • Drive MTTR/MTTD reductions and post-mortem remediation.
  • Enforce change management, release gating and production readiness reviews.
  • Partner with Bogota leadership and regional teams on governance.

Conocimientos

SRE leadership
Platform engineering
Observability
SQL Server
Communication
Troubleshooting RCA

Herramientas

ELK Stack
Application Insights
Splunk
Kibana
AppDynamics
DynaTrace

Descripción del empleo

.Job Summary

We are seeking a leadership role to run NA Commercial Insurance Production Services and Site Reliability Engineering (SRE) function out of engineering center in Bogota. This role will be responsible for the reliability, availability, and performance of critical production systems. In this role, you will bridge software engineering and operations, driving the cultural and technical transformation toward engineering-based operations. You will collaborate closely with SREs, development, testing and business teams to ensure our systems are robust and scalable effectively providing 99.9% availability. While the focus is on technical engineering, you will also mentor junior team members and share best practices with regional teams.

Key Responsibilities
Leadership & Collaboration
  • Lead a team of SREs; mentor engineers on reliability practices
  • Team building and performance evaluation of production services/SRE staff in CECC
  • Partner with development, testing and business teams to embed reliability requirements into the SDLC
  • Partner with development, infrastructure teams to ensure lower environment stability across CI footprint
  • Act as the escalation point for major incidents and production crises
    • Communicate effectively with business and operations partners, especially during critical system outages, client escalations
Reliability & Availability
  • Partner with technology and business partners to define and own SLAs, SLOs, SLIs, and error budgets across CI portfolio
  • Lead incident response, blameless post-mortems, and remediation follow-through
  • Drive reduction in MTTR and MTTD across the production estate
Process & Governance
  • Enforce change management, release gating, and production readiness reviews
    • Establish on-call practices, runbooks, and operational playbooks
    • Report on production health metrics to senior stakeholders
    • Partner with Bogota leadership team in providing oversight into AMS MSM vendor engagement. This may include weekly/monthly governance calls, site visits, performance evaluation etc.
Platform Engineering
  • Design and implement observability frameworks (metrics, logs, traces)
  • Build self-healing automation to reduce toil and manual intervention
  • Own capacity planning, performance baselining, and scalability initiatives
Skills & Experience
Required:
  • 10+ years in SRE, platform engineering, or production operations
  • 5+ years in a lead or senior individual contributor capacity
  • Strong reasoning, analytical thinking and troubleshooting skills for applications, including RCA & memory debugging.
  • Experience with observability and monitoring tools (e.g., ELK Stack, Application Insights, Splunk, Kibana, AppDynamics, DynaTrace).
  • Basic/intermediate knowledge of databases such as MS SQL Server.
  • Strong communication skills and ability to work under pressure.
Nice to Have:
  • Experience with AI technologies (e.g. Claude) and their application in enhancing system reliability and performance.
  • Experience with application production support/SRE management
  • Background in regulated or high-compliance industries.
  • Familiarity with chaos engineering, performance optimization or fault injection.
  • Familiarity with Azure cloud infrastructure and services (e.g. PaaS and identity management such as Active Directory, Azure AD).
Soft Skills:
  • Excellent verbal and written communication skills; must have strong experience in working with senior technology and business stakeholders
  • Proactive, detail-oriented, and able to handle production-critical issues.
  • Collaborative mindset and willingness to mentor others.
What Success Looks Like
  • Rapid, effective resolution of incidents and performance issues.
  • High uptime and reliability for all critical applications.
  • Continuous improvement in automation and operational efficiency.
  • A culture of reliability and technical excellence within the team.
Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Production Services Lead
Production Services Lead

Chubb • Colombia

Presencial
COP 20.000.000 - 32.000.000
Service Reliability Engineer
Service Reliability Engineer

1083 Amadeus IT Group Colombia, S.A.S. • Colombia

Presencial
COP 182.089.661 - 254.925.525
Competitive remuneration
Vacation and holiday paid time off
Health insurances
+3
Senior Site Reliability Engineer (SRE) | Bogotá
Senior Site Reliability Engineer (SRE) | Bogotá

remoti • Bogotá

Presencial
COP 300.060.000 - 433.420.000
Site Reliability Engineer
Site Reliability Engineer

DCT • Bogotá

Presencial
COP 156.225.589 - 234.338.384
Career Growth & Mentorship
Flexible Work Environment
Generative & Collaborative Culture
Site Reliability Engineer
Site Reliability Engineer

T-mapp Jobs • Bogotá

Presencial
COP 120.000.000 - 180.000.000
Hybrid work model in Bogotá
Health benefits
Learning & development programs
+1
SRE & Production Services Lead for 99.9% Reliability
SRE & Production Services Lead for 99.9% Reliability

Chubb Insurance Hong Kong • Colombia

Presencial
COP 180.000.000 - 300.000.000
Platform Operations Engineer (SRE) Colombia / Mexico
Platform Operations Engineer (SRE) Colombia / Mexico

Forte Group • Colombia

Presencial
COP 226.714.528 - 302.286.038
Systems Reliability Engineering Senior Manager
Systems Reliability Engineering Senior Manager

Scotiabank • Bogotá

Presencial
COP 180.000.000 - 300.000.000
Senior SRE & Production Services Lead
Senior SRE & Production Services Lead

Chubb • Colombia

Presencial
COP 20.000.000 - 32.000.000
Remote SRE & Production Support Engineer
Remote SRE & Production Support Engineer

Lancesoft • Bogotá ciudad

Presencial
COP 136.745.000 - 208.374.000