SRE Lead

Chubb Ltd.

Malaysia

On-site

MYR 240,000 - 420,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Chubb Ltd. is seeking an experienced Site Reliability Engineering leader in Malaysia to shape and own the reliability strategy across production systems.

You will bridge software development and operations, drive multi-year roadmaps, and mentor a high-performing SRE team prioritising automation over toil. You will champion observability, SLO/SLI governance, and data-driven reliability decisions, collaborating with engineering, product, and leadership to ensure production readiness and resilient

Qualifications

  • 8+ years in software engineering or platform reliability.
  • 3+ years in people leadership and mentoring teams.
  • Strong hands-on coding in Python or Go for automation.
  • Experience leading enterprise-scale SRE or platform reliability.
  • Knowledge of SLOs/SLIs, error budgets, and incident response.
  • Familiarity with observability stacks and CI/CD pipelines.
  • ITIL Foundation is desirable.

Responsibilities

  • Lead the SRE function and define multi-year reliability strategy.
  • Hire, onboard, and develop the SRE team and succession plans.
  • Own incident management, severity handling, and postmortems.
  • Champion blameless RCA and systemic fixes.
  • Establish and govern SLOs/SLIs and error budgets.
  • Drive resilience, chaos engineering, and production readiness.
  • Reduce toil via automation and improved runbooks.
  • Oversee observability and alerting standards across the stack.
  • Partner with engineering to embed reliability in design and delivery.

Skills

Software engineering
People leadership
Python/Go coding
Observability/Monitoring
Incident management
SDLC/DevOps
ITIL (desirable)

Education

Bachelor's in CS/IT/Eng

Tools

Dynatrace
Grafana
Prometheus
Splunk/ELK
Azure Monitor/Log Analytics
OpenTelemetry
Kubernetes/Docker
Terraform/Bicep
ServiceNow

Job description

  • Lead the Site Reliability Engineering function to define and drive the organisation’s reliability engineering strategy — bridging software development and operations through engineering discipline, not manual process.
  • Own the end-to-end reliability posture of production systems: define SLO/SLI frameworks, govern error budgets, and enforce production-readiness standards to protect business continuity.
  • Build, mentor, and scale a high-performing SRE team that prioritises engineering over toil — automating manual work, embedding reliability into the SDLC, and driving down mean time to recovery through systematic improvement.
  • Champion observability-led engineering through full-stack Dynatrace adoption, AIOps integration, and data-driven reliability decision‑making at every layer of the stack.
  • Serve as the primary reliability engineering partner to development and platform leadership, shaping architecture decisions, release policies, and automation strategy.

Key Responsibilities:

  • Define and drive the SRE strategy and multi-year roadmap aligned to business priorities.
  • Lead and develop the SRE team, including hiring, onboarding, performance management, career development, and succession planning.
  • Own incident management, including severity classification, escalation, response SLAs, and leadership of major incidents.
  • Champion blameless postmortems, root cause analysis, and implementation of systemic fixes.
  • Establish and govern SLOs, SLIs, and error budgets, ensuring reliability targets are aligned to business needs.
  • Drive resilience engineering, including chaos engineering, GameDays, production readiness reviews, and failure mode analysis.
  • Reduce toil through automation, improved runbooks, and continuous operational improvement.
  • Own observability and alerting standards, including monitoring strategy, dashboards, and alert quality.
  • Partner with engineering, architecture, product, and leadership teams to embed reliability into design and delivery.
  • Represent the SRE function in senior forums and provide reporting on reliability, risk, and operational performance.
Qualifications
  • Degree in Computer Science, Software Engineering, IT, or a related technical field.
  • 8+ years’ experience in software engineering, platform reliability, or SRE, including 3+ years in people leadership.
  • Strong hands‑on coding ability in Python, Go, or similar, with experience building automation and self‑healing solutions.
  • Proven experience leading enterprise‑scale SRE or platform reliability functions.
  • Experience defining and operating RTO/RPO and SLI/SLO frameworks.
  • Strong background in observability, production readiness, error budgets, and chaos engineering.
  • Experience leading on‑call models, incident response, and executive stakeholder engagement.
  • Solid understanding of SDLC, Agile, and DevOps delivery models.
  • ITIL Foundation is desirable.

Managerial & Soft Skills:

  • Proven people leader with experience building and coaching high-performing teams.
  • Strategic thinker who can turn business priorities into reliability roadmaps.
  • Strong communicator who can explain technical risk in business terms.
  • Calm and decisive during major incidents.
  • Influential partner across engineering, product, and leadership teams.
  • Strong advocate for developer experience and sustainable on‑call practices.
  • Data-driven and able to balance reliability, speed, and cost.
  • Champions psychological safety, continuous learning, and operational excellence.

Technical Skills:

  • Expert in Dynatrace, with experience in observability, monitoring, SLOs, tracing, and log management.
  • Proficient in Grafana, Prometheus, Splunk, ELK, Azure Monitor, and Log Analytics.
  • Strong knowledge of OpenTelemetry and telemetry pipeline design.
  • Experience with ServiceNow, CI/CD tools, Kubernetes, Docker, Terraform, and Bicep.
  • Familiar with Java, .NET, databases, APIs, Kafka, and cloud platforms, especially Azure.
  • Experience with AIOps, AI-assisted triage, and automation tooling.
  • Able to support reliability engineering through scripting, auto‑remediation, and operational automation.
  • Experience in insurance or financial services.
  • Dynatrace, Azure, ITIL 4, or Google Cloud/SRE‑related certifications.
  • Experience with chaos engineering, AIOps, MLOps, and FinOps.
  • Strong analytical skills and experience working with large operational datasets.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Lead
SRE Lead

Chubblifefund • Malaysia

On-site
MYR 250,000 - 420,000
SRE Lead
SRE Lead

Chubb • Malaysia

On-site
MYR 300,000 - 420,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Career Wise • Kuala Lumpur

On-site
MYR 80,000 - 120,000
System Reliability Engineer, Consultant
System Reliability Engineer, Consultant

AIA Malaysia • Kuala Lumpur

On-site
MYR 70,000 - 110,000
High-impact team environment
Opportunities for innovation
Influence engineering culture
System Reliability Engineer, Consultant
System Reliability Engineer, Consultant

AIA Hong Kong and Macau • Kuala Lumpur

On-site
MYR 70,000 - 90,000
Team Lead - Application DevSecOps & SRE
Team Lead - Application DevSecOps & SRE

Dialog Group Berhad • Petaling Jaya

On-site
MYR 240,000 - 360,000
Engineering Manager – Platform & SRE
Engineering Manager – Platform & SRE

INSCALE • Kuala Lumpur

On-site
MYR 350,000 - 650,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AirAsia rewards • Kuala Lumpur

On-site
MYR 180,000 - 280,000
Reliability Operations Engineer (Malaysia)
Reliability Operations Engineer (Malaysia)

Serve Robotics • Penang

On-site
MYR 485,000 - 607,000
Site Reliability Engineer
Site Reliability Engineer

LAVU TECH SOLUTIONS SDN. BHD. • Petaling Jaya

On-site
MYR 180,000 - 300,000