Senior Platform Reliability Engineer

HFG Insurance Recruitment

Cyberjaya

On-site

MYR 180,000 - 300,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

HFG Insurance Recruitment is seeking a Senior Platform Reliability Engineer to maintain the stability and efficiency of our internal container platform and supporting infrastructure in Cyberjaya.

You will manage provisioning, respond to outages, automate manual tasks with Ansible and scripting, develop Helm charts, and design Grafana/Dynatrace dashboards to track SLOs, SLAs, and performance, while ensuring compliance.

Qualifications

  • Bachelor's or Master's degree in Computer Science or related field.
  • 5–7 years IT experience; 3–5 years as Platform Reliability/SRE with container platforms (Tanzu or Kubernetes).
  • At least 3 years automation with Ansible; Python and Bash scripting experience.
  • 3 years developing Helm charts and Helm repositories.
  • Minimum 3 years managing NSX-T with Tanzu suite products.
  • Certifications: CKA/CKAD/CKS.
  • Experience in high-demand, fast-paced environments.
  • Strong monitoring design with Grafana and Dynatrace.

Responsibilities

  • Maintain stability and reliability of internal container platform and infra.
  • Respond to outages; 24/7 on-call rotation; manage production incidents.
  • Automate manual processes; implement automation to reduce errors.
  • Deploy product updates; keep platform vulnerability-free.
  • Design monitoring dashboards for SLOs/SLIs/SLAs.

Skills

Kubernetes
Platform reliability
Ansible
Python
Bash
CI/CD

Education

Bachelor's or Master's in Computer Science

Tools

Bitbucket
Nexus
Jira
Confluence
Docker
Dynatrace
Grafana

Job description

Key Responsibilities:

  • As a Senior Platform Reliability Engineer, you will play a key role in maintaining the stability, reliability, and efficiency of the organization’s internal container platform and its supporting infrastructure. Your responsibilities will include core operational tasks such as resource provisioning and management, responding to platform and application outages, and monitoring.
  • This includes proactively identifying and resolving reliability issues, analysing product dependencies, pinpointing performance bottlenecks, and implementing optimization strategies to enhance platform availability and cost efficiency.
  • In this role, you will participate in a 24/7 on-call rotation, promptly addressing alerts from the global monitoring team and resolving production incidents to maintain platform and application uptime. Additionally, you will regularly review team workflows to identify manual processes and implement automation solutions that reduce effort and minimize human error.
  • Regularly deploy product updates as required to keep the platform vulnerability-free.
  • Work with open-source technologies, CI/CD, SCM tools as necessary, and source control such as Bitbucket, implement organization containers (e.g., Docker and Kubernetes). Stay current with industry trends and propose new ways for the business to improve.
  • Take accountability in considering business and regulatory compliance risks and take appropriate steps to mitigate the risks.
  • Maintain awareness of industry trends on regulatory compliance, emerging threats and technologies to understand the risk and better safeguard the company.
  • Highlight any potential concerns/risks and proactively share best risk management practices.

We are looking for people with:

  • Bachelor's or Master's degree in Computer Science or a related field.
  • Minimum of 5 to 7 years of overall experience in IT, with at least 3 to 5 years of hands-on experience as a Platform Reliability Engineer or Site Reliability Engineer, specifically managing container orchestration platforms such as Tanzu Application Service, Tanzu Kubernetes Grid Integrated Edition, or other Kubernetes-based platforms.
  • At least 3 years of experience in automation using tools like Ansible and scripting languages such as Python and Bash.
  • 3 years of experience in developing and maintaining Helm charts and Helm repositories.
  • Minimum of 3 years of experience managing NSX-T solutions and integrating them with Tanzu suite products.
  • Possession of one or more of the following certifications:
  • a) Certified Kubernetes Administrator (CKA)
  • b) Certified Kubernetes Application Developer (CKAD)
  • c) Certified Kubernetes Security Specialist (CKS)
  • 3-5 years of experience working in high-demand, fast-paced environments.
  • Strong expertise in platform reliability principles, including scalability, performance optimization, and enterprise platform architecture.
  • Proficiency in designing monitoring dashboards using Grafana and Dynatrace to track SLOs, SLIs, and SLAs of the platform.
  • Solid understanding of DevOps pipelines and automation tools such as Bamboo, Ansible, Bitbucket, Nexus, Jira and Confluence.
  • Strong technical and business acumen with the ability to collaborate across multiple technical teams.
  • Proven experience in diagnosing and resolving infrastructure and networking issues.
  • Extensive experience in CI/CD environments, with a deep understanding of change and version control processes.
  • Hands-on experience with platform upgrades, patching, and buildpack management.
  • Ability to troubleshoot complex network-related problems.
  • Passion for continuous learning and evaluating emerging technologies, with a commitment to knowledge sharing within the team.
  • Ability to document Standard Operating Procedures (SOPs) and contribute to internal knowledge bases.
  • Strong collaboration skills with the ability to work across various stakeholder groups at an organizational level.
  • Excellent communication skills to engage with stakeholders and domain experts in designing and operating enterprise-wide solutions.
  • Self-motivated, disciplined, and proactive with a strong sense of ownership and urgency.
  • High level of integrity, takes accountability of work, and maintains a good attitude toward teamwork.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Reliability Engineer
Senior Platform Reliability Engineer

HFG (Hong Kong) Limited • Cyberjaya

On-site
MYR 180,000 - 300,000
Senior Platform Reliability Engineer Tanzu & Kubernetes
Senior Platform Reliability Engineer Tanzu & Kubernetes

Great Eastern • Cyberjaya

On-site
MYR 90,000 - 130,000
Tanzu engineer
Tanzu engineer

Encora Inc. • Kuala Lumpur

On-site
MYR 120,000 - 240,000
Senior Engineer, Platform Infrastructure (Containers & Virtualization)
Senior Engineer, Platform Infrastructure (Containers & Virtualization)

Singtel • Kuala Lumpur

On-site
Confidential
Senior Engineer, Platform Infrastructure (Containers & Virtualization)
Senior Engineer, Platform Infrastructure (Containers & Virtualization)

Singtel Group • Kuala Lumpur

On-site
MYR 120,000 - 200,000
Senior Platform Reliability Engineer - Kubernetes
Senior Platform Reliability Engineer - Kubernetes

HFG Insurance Recruitment • Cyberjaya

On-site
MYR 180,000 - 300,000
Tanzu engineer
Tanzu engineer

Linuxconfig • Kuala Lumpur

On-site
MYR 180,000 - 240,000
Senior Linux Platform Engineer
Senior Linux Platform Engineer

CLOUDENGINE DIGITAL SDN. BHD. • Kuala Lumpur

On-site
MYR 180,000 - 240,000
Senior Platform Engineer - DevOps
Senior Platform Engineer - DevOps

Endava • Kuala Lumpur

On-site
MYR 180,000 - 240,000
Share plan
Global career opportunities
Hybrid work hours
+1
OpenShift Engineer
OpenShift Engineer

Linuxconfig • Kuala Lumpur

On-site
MYR 120,000 - 160,000