Production Services Lead

Chubb

Colombia

On-site

COP 180,000,000 - 280,000,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Chubb is seeking a leadership role to run the NA Commercial Insurance Production Services and Site Reliability Engineering (SRE) function from the engineering center in Bogota. You will ensure 99.9% availability by aligning software engineering with IT operations and mentoring junior engineers across regional teams.

The role emphasizes cross-functional collaboration with development, testing, and business units to embed reliability in the SDLC and drive architectural improvements for scalable,

Qualifications

  • 10+ years in SRE, platform engineering, or production operations.
  • 5+ years in a lead or senior individual contributor capacity.
  • Strong reasoning, analytical thinking and troubleshooting skills for applications, including RCA & memory debugging.
  • Experience with observability and monitoring tools (ELK Stack, Application Insights, Splunk, Kibana, AppDynamics, DynaTrace).
  • Basic/intermediate knowledge of databases such as MS SQL Server.
  • Strong communication skills and ability to work under pressure.

Responsibilities

  • Lead a team of SREs; mentor engineers on reliability practices.
  • Team building and performance evaluation of production services/SRE staff in CECC.
  • Partner with development, testing and business teams to embed reliability requirements into the SDLC.
  • Partner with development, infrastructure teams to ensure lower environment stability across CI footprint.
  • Act as the escalation point for major incidents and production crises.
  • Define and own SLAs, SLOs, SLIs, and error budgets across CI portfolio.
  • Lead incident response, blameless post-mortems, and remediation follow-through.
  • Drive reduction in MTTR and MTTD across the production estate.
  • Enforce change management, release gating, and production readiness reviews.
  • Establish on-call practices, runbooks, and operational playbooks.
  • Report on production health metrics to senior stakeholders.
  • Partner with Bogota leadership team in overseeing AMS MSM vendor engagement.
  • Design and implement observability frameworks (metrics, logs, traces).
  • Build self-healing automation to reduce toil and manual intervention.
  • Own capacity planning, performance baselining, and scalability initiatives.

Skills

SRE
Platform engineering
Production operations
Observability
Troubleshooting
Memory debugging

Tools

ELK Stack
Application Insights
Splunk
Kibana
AppDynamics
DynaTrace

Job description

.

Job Summary

We are seeking a leadership role to run NA Commercial Insurance Production Services and Site Reliability Engineering (SRE) function out of engineering center in Bogota. This role will be responsible for the reliability, availability, and performance of critical production systems. In this role, you will bridge software engineering and operations, driving the cultural and technical transformation toward engineering-based operations. You will collaborate closely with SREs, development, testing and business teams to ensure our systems are robust and scalable effectively providing 99.9% availability. While the focus is on technical engineering, you will also mentor junior team members and share best practices with regional teams.

Key Responsibilities
Leadership & Collaboration
  • Lead a team of SREs; mentor engineers on reliability practices
  • Team building and performance evaluation of production services/SRE staff in CECC
  • Partner with development, testing and business teams to embed reliability requirements into the SDLC
  • Partner with development, infrastructure teams to ensure lower environment stability across CI footprint
  • Act as the escalation point for major incidents and production crises
    • Communicate effectively with business and operations partners, especially during critical system outages, client escalations
Reliability & Availability
  • Partner with technology and business partners to define and own SLAs, SLOs, SLIs, and error budgets across CI portfolio
  • Lead incident response, blameless post-mortems, and remediation follow-through
  • Drive reduction in MTTR and MTTD across the production estate
Process & Governance
  • Enforce change management, release gating, and production readiness reviews
    • Establish on-call practices, runbooks, and operational playbooks
    • Report on production health metrics to senior stakeholders
    • Partner with Bogota leadership team in providing oversight into AMS MSM vendor engagement. This may include weekly/monthly governance calls, site visits, performance evaluation etc.
Platform Engineering
  • Design and implement observability frameworks (metrics, logs, traces)
  • Build self-healing automation to reduce toil and manual intervention
  • Own capacity planning, performance baselining, and scalability initiatives
Skills & Experience
Required:
  • 10+ years in SRE, platform engineering, or production operations
  • 5+ years in a lead or senior individual contributor capacity
  • Strong reasoning, analytical thinking and troubleshooting skills for applications, including RCA & memory debugging.
  • Experience with observability and monitoring tools (e.g., ELK Stack, Application Insights, Splunk, Kibana, AppDynamics, DynaTrace).
  • Basic/intermediate knowledge of databases such as MS SQL Server.
  • Strong communication skills and ability to work under pressure.
Nice to Have:
  • Experience with AI technologies (e.g. Claude) and their application in enhancing system reliability and performance.
  • Experience with application production support/SRE management
  • Background in regulated or high-compliance industries.
  • Familiarity with chaos engineering, performance optimization or fault injection.
  • Familiarity with Azure cloud infrastructure and services (e.g. PaaS and identity management such as Active Directory, Azure AD).
Soft Skills:
  • Excellent verbal and written communication skills; must have strong experience in working with senior technology and business stakeholders
  • Proactive, detail-oriented, and able to handle production-critical issues.
  • Collaborative mindset and willingness to mentor others.
What Success Looks Like
  • Rapid, effective resolution of incidents and performance issues.
  • High uptime and reliability for all critical applications.
  • Continuous improvement in automation and operational efficiency.
  • A culture of reliability and technical excellence within the team.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Service Reliability Engineer
Service Reliability Engineer

1083 Amadeus IT Group Colombia, S.A.S. • Colombia

On-site
COP 182,089,661 - 254,925,525
Competitive remuneration
Vacation and holiday paid time off
Health insurances
+3
Sr. SRE Lead
Sr. SRE Lead

Chubb • Colombia

On-site
COP 120,000,000 - 180,000,000
Site Reliability Engineer
Site Reliability Engineer

T-mapp Jobs • Bogotá

On-site
COP 120,000,000 - 180,000,000
Hybrid work model in Bogotá
Health benefits
Learning & development programs
+1
Sr. SRE Lead
Sr. SRE Lead

Chubb • Norte

Hybrid
COP 287,954,000 - 447,928,000
Site Reliability Engineer
Site Reliability Engineer

DCT • Bogotá

On-site
COP 156,225,589 - 234,338,384
Career Growth & Mentorship
Flexible Work Environment
Generative & Collaborative Culture
Senior SRE & Production Services Lead
Senior SRE & Production Services Lead

Chubb • Colombia

On-site
COP 180,000,000 - 280,000,000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

FashionUnited Group • Bogotá ciudad

On-site
COP 144,000,000 - 216,000,000
Platform Operations Engineer (SRE) Colombia / Mexico
Platform Operations Engineer (SRE) Colombia / Mexico

Forte Group • Colombia

On-site
COP 226,714,528 - 302,286,038
Systems Reliability Engineering Senior Manager
Systems Reliability Engineering Senior Manager

Scotiabank • Bogotá

On-site
COP 180,000,000 - 300,000,000
Hybrid SRE: Reliability, Automation & Observability
Hybrid SRE: Reliability, Automation & Observability

T-mapp Jobs • Bogotá

Hybrid
COP 120,000,000 - 180,000,000
Hybrid work model in Bogotá
Health benefits
Learning & development programs
+1