Senior / Lead Platform Reliability Engineer (PRE)

HFG Insurance Recruitment

Cyberjaya

On-site

MYR 180,000 - 300,000

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

HFG Insurance Recruitment is seeking a Senior Lead Platform Reliability Engineer to join an enterprise infra team. You will engineer, operate and continuously improve a highly available internal container platform, using Tanzu, Kubernetes and CI/CD pipelines.

You will lead automation, troubleshooting and on-call incidents, while coordinating upgrades and platform delivery across cross-functional teams.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, IT or related field.
  • Hands-on with Tanzu TKGI or Kubernetes-based platforms.
  • Experience in Ansible, Python and Bash scripting.
  • Developing and maintaining Helm charts and repositories.
  • Experience with NSX-T and Tanzu/Kubernetes integration.
  • Strong Kubernetes, Docker and container tech knowledge.
  • CI/CD pipelines experience and DevOps practices.
  • Experience with platform upgrades, patching and buildpack management.
  • Strong troubleshooting across infra, networking and container platforms.
  • Monitoring with Grafana or Dynatrace for SLOs/SLIs/SLAs.

Responsibilities

  • Engineer and maintain enterprise-grade container platforms with focus on reliability and performance.
  • Work with Broadcom VMware Tanzu and Kubernetes-based orchestration.
  • Handle capacity planning, monitoring and reliability improvements.
  • Provide L2/L3 production support and on-call incident response.
  • Lead or support major incident management and remediation.
  • Design automation via Ansible, Python, and Bash to reduce manual toil.
  • Develop and maintain Helm charts and Helm repositories.
  • Collaborate with DevOps tools like Docker, CI/CD and source control.
  • Manage platform upgrades, patches and buildpack management.
  • Monitor platform reliability using Grafana/Dynatrace, targeting SLOs/SLIs/SLAs.
  • Coordinate with multiple teams to deliver reliable enterprise solutions.

Skills

Kubernetes/Tanzu
Automation
CI/CD
Networking
Platform reliability
Ansible
Python
Bash
Helm
NSX-T
Grafana/Dynatrace
Docker
Bitbucket/Jira/Confluence

Education

Bachelor's or Master's in CS/IT

Tools

TAS/TKGI/Kubernetes platforms
Ansible
Python
Helm charts
NSX-T
Grafana
Dynatrace
Docker
CI/CD tooling
Bitbucket/Nexus/Jira/Confluence

Job description

Hiring: Senior / Lead Platform Reliability Engineer (PRE)

We are looking for experienced Platform Reliability Engineers to join an enterprise infrastructure team responsible for engineering, operating and continuously improving a highly available internal container platform.

This is an excellent opportunity for engineers with strong Tanzu / Kubernetes, automation, CI/CD, networking and platform reliability experience who enjoy working on large-scale enterprise environments.

Key Responsibilities
  • Engineer, operate and maintain enterprise-grade container platforms and supporting infrastructure, with a strong focus on reliability, resiliency, security and performance.
  • Work extensively with Broadcom VMware Tanzu and Kubernetes-based container orchestration platforms.
  • Perform platform resource provisioning, capacity planning, monitoring, performance optimization and reliability improvements.
  • Provide L2/L3 production support, including troubleshooting complex platform, infrastructure, application and networking issues.
  • Participate in a 24/7 on-call rotation and respond to critical production alerts and incidents.
  • Lead or support major incident management, including troubleshooting, vendor coordination, immediate remediation, root cause analysis and long-term corrective actions.
  • Design and implement automation using Ansible, Python and Bash to reduce manual processes and operational errors.
  • Develop and maintain Helm charts and Helm repositories.
  • Work with Docker, Kubernetes, CI/CD and source-control technologies, including Bitbucket and related DevOps tooling.
  • Manage platform upgrades, patching, product updates and buildpack management.
  • Monitor and optimize platform reliability using Grafana and Dynatrace, with a focus on SLOs, SLIs and SLAs.
  • Troubleshoot and resolve complex infrastructure, networking and container platform issues.
  • Review security advisories and ensure timely remediation and updates across the container platform.
  • Work closely with application, infrastructure, security, network and other technical teams to deliver reliable enterprise solutions.
  • For the Lead level, provide technical leadership to an existing PRE team, manage on-call resources, coordinate platform upgrades and deployment activities, and oversee ServiceNow/Jira queues and SLA delivery.
Technical Requirements
  • Bachelor's or Master's degree in Computer Science, IT or a related discipline.
  • Strong hands‑on experience with Tanzu Application Service (TAS), Tanzu Kubernetes Grid Integrated Edition (TKGI), or Kubernetes-based platforms.
  • Strong experience with Ansible, Python and Bash scripting.
  • Hands‑on experience developing and maintaining Helm charts and Helm repositories.
  • Experience with NSX‑T and integration with Tanzu/Kubernetes environments.
  • Strong understanding of Kubernetes, Docker and container technologies.
  • Experience with CI/CD pipelines, SCM and DevOps practices.
  • Experience with platform upgrades, patching and buildpack management.
  • Strong troubleshooting capabilities across infrastructure, networking and container platforms.
  • Experience with Grafana and/or Dynatrace, including monitoring SLOs, SLIs and SLAs.
  • Familiarity with Bamboo, Bitbucket, Nexus, Jira and Confluence.
  • Strong understanding of reliability engineering, scalability, performance optimization and enterprise platform architecture.
  • Excellent stakeholder management, communication and documentation skills.
Senior Platform Reliability Engineer
  • 5–7 years of overall IT experience.
  • 3–5 years of hands‑on experience in Platform Reliability Engineering or Site Reliability Engineering.
  • 3+ years of automation experience using Ansible, Python and Bash.
  • 3+ years of Helm experience.
  • 3+ years of NSX‑T experience.
  • Experience operating in high‑demand, fast‑paced production environments.
Lead Platform Reliability Engineer
  • 7–10+ years of hands‑on experience with container orchestration platforms.
  • 5+ years of automation experience using Ansible, Python and Bash.
  • 5+ years of Helm experience.
  • 3+ years of NSX‑T and Tanzu integration experience.
  • Strong experience in enterprise platform architecture and reliability engineering.
  • Proven experience leading technical teams and managing production operations.
  • Experience owning L3 support, major incidents, platform upgrades and operational delivery.
Certifications

One or more of the following is highly desirable:

  • Certified Kubernetes Security Specialist (CKS)
Ideal Candidate

We are looking for someone who is technically hands‑on, proactive and comfortable taking ownership of production platforms. You should enjoy solving complex infrastructure and networking problems, automating repetitive processes and continuously improving platform reliability.

Candidates with strong Tanzu/Kubernetes, NSX‑T, Ansible, Helm and CI/CD experience are encouraged to apply.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Reliability Engineer
Senior Platform Reliability Engineer

HFG (Hong Kong) Limited • Cyberjaya

On-site
MYR 180,000 - 300,000
Senior Platform Reliability Engineer
Senior Platform Reliability Engineer

HFG Insurance Recruitment • Putrajaya, Cyberjaya

On-site
MYR 180,000 - 240,000
Senior Platform Reliability Engineer Tanzu & Kubernetes
Senior Platform Reliability Engineer Tanzu & Kubernetes

Great Eastern • Cyberjaya

On-site
MYR 90,000 - 130,000
Senior Platform Reliability Engineer - Tanzu & Kubernetes
Senior Platform Reliability Engineer - Tanzu & Kubernetes

HFG Insurance Recruitment • Cyberjaya

On-site
MYR 180,000 - 300,000
Senior Platform Reliability Engineer: Kubernetes Automation
Senior Platform Reliability Engineer: Kubernetes Automation

HFG Insurance Recruitment • Putrajaya, Cyberjaya

On-site
MYR 180,000 - 240,000
Tanzu Engineer
Tanzu Engineer

Chemcastle Sdn Bhd • Kuala Lumpur

On-site
MYR 240,000 - 320,000
Senior Engineer, Platform Infrastructure (Containers & Virtualization)
Senior Engineer, Platform Infrastructure (Containers & Virtualization)

Singtel Group • Kuala Lumpur

On-site
MYR 120,000 - 200,000
Senior Engineer, Platform Infrastructure (Containers & Virtualization)
Senior Engineer, Platform Infrastructure (Containers & Virtualization)

Singtel • Kuala Lumpur

On-site
Confidential
Senior Platform Reliability Engineer: Container & Cloud
Senior Platform Reliability Engineer: Container & Cloud

Great Eastern • Cyberjaya

On-site
MYR 180,000 - 290,000
Senior DevOps Engineer
Senior DevOps Engineer

Involve Asia • Kuala Lumpur

On-site
MYR 180,000 - 300,000