Senior Manager, DevOps & Production Support Engineer

FWD Insurance

Hong Kong

On-site

HKD 1,200,000 - 2,000,000

Full time

6 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

FWD Group in Hong Kong is looking for a senior DevOps/SRE leader to own Production Support for Group Digital Platforms, driving reliability, incident response, and platform stability across regional markets. You will define and govern enterprise DevOps practices, including IaC, CI/CD pipelines, cloud security controls, and observability.

Lead cross-functional teams with Product, Engineering, Security, and Market stakeholders to balance scalability, resilience, and cost while delivering superior

Qualifications

  • Senior DevOps/SRE professional with 10+ years managing large-scale platforms.
  • Strong knowledge of CI/CD, IaC, observability, incident management, and automation.
  • Experience across cloud providers (AWS/Azure/GCP) and container orchestration.

Responsibilities

  • Lead DevOps and Production Support for Group Digital Platforms across regional markets.
  • Define, implement, and govern IaC, CI/CD pipelines, cloud security, and platform engineering standards.
  • Own Level 1 production support operations, including incident management and service restoration processes.
  • Drive observability strategies and continuous service improvement to minimise incidents.

Skills

DevOps leadership
SRE / Platform engineering
Cloud architecture
AWS / Azure
Kubernetes
CI/CD pipelines
Observability
Incident management

Education

Bachelor's degree in CS/Engineering

Tools

Terraform
Docker
CloudFormation
Ansible

Job description

About FWD Group

FWD Group (1828.HK) is a pan-Asian life and health insurance business that serves approximately 40 million customers across 10 markets, including BRI Life in Indonesia. FWD’s customer-led and tech-enabled approach aims to deliver innovative propositions, easy-to-understand products and a simpler insurance experience. Established in 2013, the company operates in some of the fastest-growing insurance markets in the world with a vision of changing the way people feel about insurance. FWD Group is listed on the main board of the Hong Kong Stock Exchange under the stock code 1828.

For more information, please visit www.fwd.com

PURPOSE

Own the operational reliability of Group Digital Platforms by leading the Production Support DevOps Engineering function, ensuring timely incident response, effective problem resolution, proactive monitoring, and consistent operational governance across regional markets.

KEY ACCOUNTAIBILITIES
  • Lead the DevOps and Production Support function for Group Digital Platforms, with accountability for platform reliability, availability, operational resilience, and service excellence across regional markets.
  • Define, implement, and govern enterprise DevOps practices, including infrastructure automation, Infrastructure-as-Code (IaC), CI/CD pipelines, cloud security controls, and platform engineering standards.
  • Own Level 1 production support operations, ensuring effective incident management, major incident response, problem management, escalation governance, and service restoration processes.
  • Establish and maintain highly resilient cloud platforms across AWS and Azure, incorporating self-healing capabilities, auto-scaling, disaster recovery, high availability, and business continuity requirements.
  • Drive operational excellence through automation, reducing manual intervention across infrastructure provisioning, deployments, monitoring, alerting, and routine support activities.
  • Lead root cause analysis and continuous service improvement initiatives to minimise recurring incidents and improve platform stability, reliability, and customer experience.
  • Define and implement comprehensive observability strategies, including monitoring, logging, alerting, tracing, and business transaction visibility across platform and application services.
  • Partner closely with Product Owners, Engineering, Architecture, Security, and Market teams to ensure production services operate efficiently with minimal customer impact and optimal service performance.
  • Govern release management processes, ensuring production readiness, deployment quality, operational risk assessment, rollback preparedness, and post-release monitoring.
  • Drive cloud architecture design, implementation, and optimisation, balancing scalability, resilience, security, performance, and cost efficiency across digital platforms.
  • Provide technical leadership and guidance on platform architecture, cloud-native solutions, operational best practices, and reliability engineering principles.
  • Influence technology and operational decisions across Group and Markets, ensuring alignment with enterprise standards, operational objectives, and platform strategy.
  • Establish and maintain DevOps engineering standards, operational procedures, runbooks, support playbooks, and technical documentation to ensure consistency and operational maturity.
  • Lead cross-functional collaboration during critical incidents, platform upgrades, service transitions, and large-scale transformation initiatives.
  • Manage operational priorities, resource allocation, delivery commitments, and stakeholder expectations across multiple markets and business functions.
  • Build internal platform capabilities by developing automation frameworks, engineering standards, reusable cloud services, and operational best practices.
  • Collaborate with technology partners, cloud providers, and strategic vendors to continuously improve platform capabilities, innovation, service reliability, and operational efficiency.
  • Ensure compliance with security, risk, governance, regulatory, and audit requirements while maintaining delivery agility and operational effectiveness.
  • Continuously optimise cloud utilisation, infrastructure performance, and operational costs through FinOps practices, capacity planning, and proactive resource management.
QUALIFICATIONS / EXPERIENCE
  • Bachelor’s degree in computer science, Engineering, Information Technology, or a related discipline, or equivalent practical experience.
  • Hands-on experience in production support, site reliability engineering, DevOps, infrastructure operations, or platform engineering within a large-scale technology environment.
  • Exposure to software development lifecycle (SDLC) practices and proficiency in one or more programming or scripting languages (e.g., Java, Python, JavaScript, PowerShell, Bash) is preferred.
  • Experience working in fast-paced, high-growth, or digitally driven environments is advantageous.
  • Strong understanding of modern DevOps practices, including CI/CD, Infrastructure as Code (IaC), observability, monitoring, incident management, problem management, and operational automation.
  • Proven experience supporting critical production systems with a focus on service reliability, availability, performance, and operational resilience.
  • Familiarity with cloud platforms (AWS, Azure, or GCP), containerisation technologies, and enterprise operational support frameworks is preferred.
  • Strong analytical, troubleshooting, stakeholder management, and communication skills with the ability to operate effectively during major incidents and service disruptions.
KNOWLEDGE & TECHNICAL SKILLS
  • Minimum 10-12 years of experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, Production Operations, or related disciplines, of which at least 5 years in a senior technical leadership or management capacity supporting large-scale digital platforms.
  • Demonstrated expertise in designing and governing enterprise-grade cloud infrastructure, container platforms, CI/CD pipelines, and production support operating models across multiple teams and markets
  • Strong knowledge of containerisation and orchestration technologies, including Docker and Kubernetes, with experience driving platform scalability, reliability, security, and operational excellence
  • Extensive experience with cloud platforms such as AWS and/or Microsoft Azure, including the successful delivery and operation of large-scale, business-critical systems in highly available environments.
  • Deep understanding of infrastructure architecture, including networking, load balancing, CDN, DNS, API gateway, security controls, disaster recovery, and high-availability design patterns.
  • Strong knowledge of modern DevOps practices, including Infrastructure as Code (Terraform, CloudFormation, Bicep), configuration management, CI/CD automation, release management, and platform engineering.
  • Proven experience establishing observability strategies and operational monitoring frameworks leveraging tools such as ElasticCloud, Dynatrace, New Relic, Elastic, Grafana, Prometheus, or equivalent enterprise platforms.
  • Strong knowledge of incident management, problem management, root cause analysis, service reliability engineering, operational resilience, and continuous service improvement practices.
  • Solid understanding of Linux/Unix platforms, networking fundamentals, cloud-native architectures, and application runtime environments.
  • Experience with software engineering practices and proficiency in one or more programming or scripting languages such as Python, Node.js, Java, PowerShell, or Bash for automation and operational efficiency.
  • Strong understanding of security, risk management, compliance, governance, and operational controls within enterprise and regulated environments.
  • Excellent stakeholder management, communication, and influencing skills, with the ability to translate complex technical concepts into clear business outcomes and executive-level recommendations.
  • Proven ability to lead cross-functional teams during major incidents, critical service disruptions, large-scale platform transformations, and complex production support environments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Manager, DevOps & Production Support Engineer
Senior Manager, DevOps & Production Support Engineer

FWD Group Management Holdings Limited • Hong Kong

On-site
HKD 900,000 - 1,800,000
Assistant Manager, Cloud Management
Assistant Manager, Cloud Management

fwd • Hong Kong

On-site
HKD 900,000 - 1,200,000
Assistant Manager, Cloud Management
Assistant Manager, Cloud Management

FWD Insurance • Hong Kong

On-site
HKD 900,000 - 1,300,000
System Analyst
System Analyst

fwd • Hong Kong

On-site
HKD 900,000 - 1,200,000
Assistant Manager, Cloud Management
Assistant Manager, Cloud Management

FWD Group Management Holdings Limited • Hong Kong

Hybrid
HKD 1,000,000 - 1,600,000
System Analyst
System Analyst

FWD Insurance • Hong Kong

On-site
HKD 600,000 - 840,000
Manager, Business Analyst
Manager, Business Analyst

FWD Group Management Holdings Limited • Hong Kong

On-site
HKD 600,000 - 900,000
System Analyst, Application Management(Frontend Distribution Platform)
System Analyst, Application Management(Frontend Distribution Platform)

FWD Group Management Holdings Limited • Hong Kong

On-site
HKD 500,000 - 800,000
System Analyst
System Analyst

FWD Group Management Holdings Limited • Hong Kong

On-site
HKD 600,000 - 900,000
Group Vice President, Global High Net Worth Operating Office
Group Vice President, Global High Net Worth Operating Office

FWD Insurance • Hong Kong

On-site
HKD 1,200,000 - 1,900,000