Sr. SRE Lead

Chubb

Norte

Hybrid

COP 287,954,000 - 447,928,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Chubb is seeking a hands-on Site Reliability Engineer (SRE) to join our LATAM/North America team. You will troubleshoot, maintain, and improve the reliability of critical applications and infrastructure across on-prem and Azure environments.

You will lead incident response, implement observability, automation, and CI/CD enhancements, and mentor junior engineers while collaborating with regional teams to share best practices. Strong Windows, IIS, .NET, and scripting skills are required.

Qualifications

  • 8+ years of hands-on SRE or application support engineering experience.
  • Strong troubleshooting for applications and servers (RCA, memory, perf).
  • Proficiency in PowerShell, Python, and .NET with automation focus.
  • Experience with Windows internals, IIS admin, and deployments.
  • Familiar with observability tools and cloud (Azure).
  • Excellent communication and ability to work under pressure.

Responsibilities

  • Hands-on engineering, troubleshooting, and performance optimization.
  • Set up and maintain Windows environments and IIS deployments.
  • Deploy and manage apps in on-prem and Azure environments.
  • Develop automation scripts to streamline deployments and monitoring.
  • Lead incident management, postmortems, and prevention measures.
  • Mentor engineers and share reliability best practices.

Skills

SRE experience
Troubleshooting
PowerShell
Python
.NET
IIS administration
Azure
Observability tools
Automation
Communication skills

Tools

AppDynamics
New Relic
OpenTelemetry
Splunk
ELK Stack
Azure Insights
DynaTrace
ScienceLogic

Job description

.Job Summary

We are looking for a highly skilled, hands-on Site Reliability Engineer (SRE) to join our LATAM/North America team. In this role, you will be directly responsible for troubleshooting, maintaining, and improving the reliability and performance of critical applications and infrastructure. You will work closely with other SREs, developers, and business teams to ensure our systems are robust, observable, and scalable. While this is primarily a technical engineering role, you will also provide mentorship and occasional guidance to junior team members and collaborate with regional teams to share best practices.

Key Responsibilities
Hands-On Engineering & Troubleshooting
  • Diagnose and resolve complex application and server issues, including root cause analysis (RCA), memory debugging, and performance optimization.
  • Perform application and server troubleshooting, including IIS administration and environment setup/deployment on Windows.
  • Write and maintain automation scripts (PowerShell, Python) to streamline deployments, monitoring, and operational tasks.
  • Develop and debug .NET applications as needed to support reliability and performance goals.
Monitoring, Observability & Analysis
  • Set up and maintain observability tools and dashboards (e.g., AppDynamics, New Relic, OpenTelemetry) to monitor application health and user experience.
  • Aggregate and analyze logs using tools like Splunk or ELK Stack to identify and resolve performance bottlenecks.
  • Define and track Critical User Journeys (CUJs) and ensure relevant telemetry is in place.
Infrastructure & Deployment
  • Set up, configure, and maintain Windows environments and server configurations.
  • Deploy and manage applications in both on-premises and cloud (Azure) environments.
  • Manage identity and access (Active Directory, Azure AD) and leverage Platform-as-a-Service (PaaS) offerings as needed.
Automation & Toil Reduction
  • Identify repetitive manual tasks and automate them to improve operational efficiency and reduce toil.
  • Advocate for and implement process improvements and tooling enhancements.
Incident Response & Reliability
  • Deep understanding of SRE principles, including SLAs, SLOs, error budgets, and reliability-focused system design.
  • Lead incident management, postmortems, and implement preventative measures.
  • Support chaos engineering and resilience testing to ensure robust recovery mechanisms.
Collaboration & Mentorship
  • Mentor engineers and champion SRE best practices, embedding a reliability-first culture and ensuring technical excellence across engineering teams.
  • Collaborate with cross-functional teams (developers, infrastructure, business) to drive reliability improvements.
  • Share knowledge and mentor junior SREs, promoting best practices and a culture of reliability.
Skills & Experience
Required:
  • 8+ years of hands-on SRE or application support engineering experience, with some exposure to team or project leadership.
  • Strong troubleshooting skills for applications and servers, including RCA, memory debugging, and performance tuning.
  • Proficiency in PowerShell scripting, Python programming, and .NET development/debugging with a focus on automation, tooling, and system integration.
  • Experience with Windows internals, IIS administration, and environment setup/deployment.
  • Deep familiarity with observability and monitoring tools (AppDynamics, DynaTrace, ScienceLogic, Azure Insights, Splunk, ELK Stack, OpenTelemetry).
  • Analytical thinking and advanced troubleshooting skills.
  • Experience with Azure cloud infrastructure and services.
  • Strong communication skills and ability to work under pressure.
Nice to Have:
  • Experience with both application and infrastructure SRE roles.
  • Background in regulated or high-compliance industries.
  • Experience with chaos engineering, performance optimization, or fault injection.
  • Knowledge of PaaS and identity management (Active Directory, Azure AD).
Soft Skills:
  • Proactive, detail-oriented, and able to handle production-critical issues.
  • Collaborative mindset and willingness to mentor others.
What Success Looks Like
  • Rapid, effective resolution of incidents and performance issues.
  • High uptime and reliability for all critical applications.
  • Continuous improvement in automation and operational efficiency.
  • A culture of reliability and technical excellence within the team.
Join Us

If you are a hands-on SRE engineer passionate about technology and reliability, and eager to make a direct impact, we would love to hear from you.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. SRE Lead
Sr. SRE Lead

Chubb • Colombia

On-site
COP 120,000,000 - 180,000,000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Joinimagine • Bogotá ciudad, Guavio

On-site
COP 9,000,000 - 13,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

N-iX • Colombia

On-site
COP 120,000,000 - 180,000,000
Flexible work format
Education reimbursement
Professional development
Production Services Lead
Production Services Lead

Chubb • Colombia

On-site
COP 180,000,000 - 280,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Michael Page Colombia • San Gil

On-site
COP 228,815,498 - 343,223,247
Crecimiento profesional a través de desafíos técnicos
Trabajo con tecnologías cloud de vanguardia
Exposición a prácticas SRE modernas
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

FashionUnited Group • Bogotá ciudad

On-site
COP 144,000,000 - 216,000,000
Service Reliability Engineer
Service Reliability Engineer

1083 Amadeus IT Group Colombia, S.A.S. • Colombia

On-site
COP 182,089,661 - 254,925,525
Competitive remuneration
Vacation and holiday paid time off
Health insurances
+3
Site Reliability Engineer ID53670
Site Reliability Engineer ID53670

AgileEngine • Metropolitana

On-site
COP 142,369,020 - 213,553,530
Professional growth: Mentorship, TechTalks, and personalized growth roadmaps.
Competitive compensation: USD-based pay with education, fitness, and team activity budgets.
Exciting projects: Modern solutions with Fortune 500 and top product companies.
+1
Senior SRE Lead: Reliability, Automation & Mentorship
Senior SRE Lead: Reliability, Automation & Mentorship

Chubb • Colombia

On-site
COP 120,000,000 - 180,000,000
Senior SRE Lead - Reliability, Automation & Observability
Senior SRE Lead - Reliability, Automation & Observability

Chubb • Norte

Hybrid
COP 287,954,000 - 447,928,000