Senior Platform Reliability Engineer

HFG (Hong Kong) Limited

Cyberjaya

On-site

MYR 180,000 - 300,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

HFG (Hong Kong) Limited in Malaysia seeks a Senior Platform Reliability Engineer to maintain the stability and performance of our internal container platform and infrastructure.

You will be on a 24/7 on-call rotation, implement automation, and design Grafana/Dynatrace dashboards to track SLOs, SLIs, and SLAs, ensuring reliability and cost efficiency across environments.

Qualifications

  • Bachelor's or Master's degree in Computer Science or a related field.
  • 5-7 years IT experience, with 3-5 years in Platform Reliability Engineer or Site Reliability Engineer roles.
  • Experience managing container orchestration platforms such as Tanzu Application Service or Tanzu Kubernetes Grid Integrated Edition, or other Kubernetes-based platforms.
  • At least 3 years of automation using Ansible and scripting in Python and Bash.
  • 3 years developing and maintaining Helm charts and Helm repositories.
  • Minimum 3 years with NSX-T and integrating with Tanzu suite products.
  • One or more of the following certifications: CKA, CKAD, or CKS.
  • 3-5 years in high-demand, fast-paced environments.
  • Strong expertise in platform reliability principles (scalability, performance, enterprise platform architecture).
  • Proficiency in designing monitoring dashboards using Grafana and Dynatrace to track SLOs, SLIs, and SLAs.

Responsibilities

  • Maintain stability, reliability, and efficiency of the internal container platform and supporting infrastructure.
  • Respond to platform and application outages and perform monitoring.
  • Proactively identify and resolve reliability issues, analyse dependencies, and pinpoint performance bottlenecks.
  • Implement optimization strategies to enhance availability and cost efficiency.
  • Participate in 24/7 on-call rotation, addressing alerts and resolving production incidents.
  • Review team workflows to identify manual processes and implement automation solutions.
  • Deploy product updates to keep the platform vulnerability‑free.
  • Work with open-source technologies, CI/CD, SCM tools, and Bitbucket.
  • Implement organization containers such as Docker and Kubernetes.
  • Take accountability for business and regulatory compliance risks and implement mitigation steps.

Skills

Container orchestration
Kubernetes
Tanzu
Ansible
Python
Bash
Helm
NSX-T
Grafana
Dynatrace
SRE principles

Education

Bachelor's or Master's degree in Computer Science or related field

Tools

Kubernetes
Tanzu Kubernetes Grid Integrated Edition
Tanzu Application Service
Docker
Bitbucket
Helm

Job description

As a Senior Platform Reliability Engineer, you will play a key role in maintaining the stability, reliability, and efficiency of the organization's internal container platform and its supporting infrastructure. You will participate in a 24/7 on-call rotation, promptly addressing alerts from the global monitoring team and resolving production incidents to maintain platform and application uptime.

Key responsibilities
  • Maintain stability, reliability, and efficiency of the organization's internal container platform and supporting infrastructure
  • Respond to platform and application outages and perform monitoring
  • Proactively identify and resolve reliability issues, analyse product dependencies, and pinpoint performance bottlenecks
  • Implement optimization strategies to enhance platform availability and cost efficiency
  • Participate in 24/7 on-call rotation, addressing alerts and resolving production incidents
  • Review team workflows to identify manual processes and implement automation solutions
  • Deploy product updates to keep the platform vulnerability‑free
  • Work with open-source technologies, CI/CD, SCM tools, and source control such as Bitbucket
  • Implement organization containers such as Docker and Kubernetes
  • Take accountability in considering business and regulatory compliance risks and implement appropriate mitigation steps
About you
  • Bachelor's or Master's degree in Computer Science or a related field
  • Minimum of 5 to 7 years of overall experience in IT, with at least 3 to 5 years of hands‑on experience as a Platform Reliability Engineer or Site Reliability Engineer
  • Specific experience managing container orchestration platforms such as Tanzu Application Service, Tanzu Kubernetes Grid Integrated Edition, or other Kubernetes-based platforms
  • At least 3 years of experience in automation using tools like Ansible and scripting languages such as Python and Bash
  • 3 years of experience in developing and maintaining Helm charts and Helm repositories
  • Minimum of 3 years of experience managing NSX‑T solutions and integrating them with Tanzu suite products
  • One or more of the following certifications: Certified Kubernetes Administrator (CKA), Certified Kubernetes Application Developer (CKAD), or Certified Kubernetes Security Specialist (CKS)
  • 3–5 years of experience working in high-demand, fast‑paced environments
  • Strong expertise in platform reliability principles, including scalability, performance optimization, and enterprise platform architecture
  • Proficiency in designing monitoring dashboards using Grafana and Dynatrace to track SLOs, SLIs, and SLAs
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Reliability Engineer
Senior Platform Reliability Engineer

HFG Insurance Recruitment • Putrajaya, Cyberjaya

On-site
MYR 180,000 - 240,000
Senior / Lead Platform Reliability Engineer (PRE)
Senior / Lead Platform Reliability Engineer (PRE)

HFG Insurance Recruitment • Cyberjaya

On-site
MYR 180,000 - 300,000
Senior Platform Reliability Engineer: Kubernetes Automation
Senior Platform Reliability Engineer: Kubernetes Automation

HFG Insurance Recruitment • Putrajaya, Cyberjaya

On-site
MYR 180,000 - 240,000
Senior Platform Reliability Engineer - Tanzu & Kubernetes
Senior Platform Reliability Engineer - Tanzu & Kubernetes

HFG Insurance Recruitment • Cyberjaya

On-site
MYR 180,000 - 300,000
Senior Platform Reliability Engineer Tanzu & Kubernetes
Senior Platform Reliability Engineer Tanzu & Kubernetes

Great Eastern • Cyberjaya

On-site
MYR 90,000 - 130,000
Senior Engineer, Platform Infrastructure (Containers & Virtualization)
Senior Engineer, Platform Infrastructure (Containers & Virtualization)

Singtel Group • Kuala Lumpur

On-site
MYR 120,000 - 200,000
Senior Engineer, Platform Infrastructure (Containers & Virtualization)
Senior Engineer, Platform Infrastructure (Containers & Virtualization)

Singtel • Kuala Lumpur

On-site
Confidential
Tanzu Engineer
Tanzu Engineer

Chemcastle Sdn Bhd • Kuala Lumpur

On-site
MYR 240,000 - 320,000
Tanzu engineer
Tanzu engineer

Encora Inc. • Kuala Lumpur

On-site
MYR 120,000 - 240,000
Senior Platform Reliability Engineer: Container & Cloud
Senior Platform Reliability Engineer: Container & Cloud

Great Eastern • Cyberjaya

On-site
MYR 180,000 - 290,000