Platform Reliability Engineer - Tanza

Coforge

Kuala Lumpur

On-site

MYR 240,000 - 360,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Coforge is seeking an experienced Platform Reliability Engineer (PRE/SRE) to architect, deploy, and operate enterprise container platforms on VMware Tanzu, PCF, Kubernetes, and cloud environments. The role emphasizes reliability, performance, and scalable design.

You will own CI/CD pipelines, IaC, monitoring, incident response, and production support, collaborating with Infrastructure, Cloud, DevOps, and Security teams.

Qualifications

  • 5+ years of experience in Platform Engineering, Production Reliability Engineering (PRE), or Site Reliability Engineering (SRE).
  • 3-5+ years of hands-on experience with VMware Tanzu and/or Pivotal Cloud Foundry (PAS, PKS, PCF).
  • 4-6+ years of Kubernetes and container platform administration experience.
  • 4-6+ years supporting CI/CD pipelines and DevOps toolchains.
  • 3-5+ years of Infrastructure-as-Code and automation experience using Terraform, Ansible, Chef, or similar tools.
  • 3-5+ years of programming/scripting experience using Python, Go, Java, or equivalent languages.
  • Strong experience with monitoring, observability, incident management, and production support.
  • Hands-on experience with Docker and container lifecycle management.
  • Strong knowledge of Git-based version control and branching strategies.

Responsibilities

  • Manage, administer, and support VMware Tanzu and/or Pivotal Cloud Foundry (PCF/PAS/PKS) platforms.
  • Deploy, maintain, and troubleshoot Kubernetes clusters and containerized applications.
  • Ensure platform reliability, availability, scalability, and performance across production environments.
  • Design, build, and maintain CI/CD pipelines using tools such as Jenkins, Bamboo, GitHub, Bitbucket, and Nexus.
  • Develop Infrastructure-as-Code (IaC) solutions using Terraform, Ansible, Chef, or similar automation tools.
  • Create automation scripts and utilities using Python, Go, Java, or other programming languages.
  • Implement and maintain monitoring, logging, and observability solutions using Prometheus, Grafana, ELK, and related tools.
  • Support cloud-based workloads across AWS, Azure, VMware, or hybrid cloud environments.
  • Lead incident response activities, root cause analysis (RCA), and postmortem reviews to improve platform stability.
  • Drive operational excellence through automation, proactive monitoring, and continuous reliability improvements.
  • Manage Docker container lifecycle and support container platform operations.
  • Collaborate with Infrastructure, Cloud, DevOps, Security, and Application teams to deliver platform enhancements and resolve complex issues.
  • Mentor junior engineers and provide technical leadership for platform operations and reliability initiatives.

Job description

Job Title: Platform Reliability Engineer (PRE) - Tanza

Job Summary:

We are seeking an experienced Platform Reliability Engineer (PRE) / Site Reliability Engineer (SRE) with strong expertise in VMware Tanzu, Pivotal Cloud Foundry (PCF), Kubernetes, cloud platforms, and DevOps automation. The ideal candidate will be responsible for ensuring the reliability, scalability, performance, and availability of enterprise container platforms and cloud-native applications. This role requires hands‑on experience in platform engineering, infrastructure automation, CI/CD, monitoring, and production support.

Key Responsibilities
  • Manage, administer, and support VMware Tanzu and/or Pivotal Cloud Foundry (PCF/PAS/PKS) platforms.
  • Deploy, maintain, and troubleshoot Kubernetes clusters and containerized applications.
  • Ensure platform reliability, availability, scalability, and performance across production environments.
  • Design, build, and maintain CI/CD pipelines using tools such as Jenkins, Bamboo, GitHub, Bitbucket, and Nexus.
  • Develop Infrastructure-as-Code (IaC) solutions using Terraform, Ansible, Chef, or similar automation tools.
  • Create automation scripts and utilities using Python, Go, Java, or other programming languages.
  • Implement and maintain monitoring, logging, and observability solutions using Prometheus, Grafana, ELK, and related tools.
  • Support cloud-based workloads across AWS, Azure, VMware, or hybrid cloud environments.
  • Lead incident response activities, root cause analysis (RCA), and postmortem reviews to improve platform stability.
  • Drive operational excellence through automation, proactive monitoring, and continuous reliability improvements.
  • Manage Docker container lifecycle and support container platform operations.
  • Collaborate with Infrastructure, Cloud, DevOps, Security, and Application teams to deliver platform enhancements and resolve complex issues.
  • Mentor junior engineers and provide technical leadership for platform operations and reliability initiatives.
Required Qualifications
  • 5+ years of experience in Platform Engineering, Production Reliability Engineering (PRE), or Site Reliability Engineering (SRE).
  • 3-5+ years of hands‑on experience with VMware Tanzu and/or Pivotal Cloud Foundry (PAS, PKS, PCF).
  • 4-6+ years of Kubernetes and container platform administration experience.
  • 4-6+ years supporting CI/CD pipelines and DevOps toolchains.
  • 3-5+ years of Infrastructure-as-Code and automation experience using Terraform, Ansible, Chef, or similar tools.
  • 3-5+ years of programming/scripting experience using Python, Go, Java, or equivalent languages.
  • Strong experience with monitoring, observability, incident management, and production support.
  • Hands‑on experience with Docker and container lifecycle management.
  • Strong knowledge of Git‑based version control and branching strategies.
Preferred Qualifications
  • Experience supporting mission‑critical enterprise platforms in highly available environments.
  • Certifications in Kubernetes, VMware Tanzu, AWS, Azure, or related cloud technologies.
  • Experience leading small engineering teams and engaging with cross‑functional stakeholders.
  • Strong understanding of SRE principles, platform automation, and cloud‑native architecture.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Reliability Engineer – Kubernetes & Cloud
Platform Reliability Engineer – Kubernetes & Cloud

Coforge • Kuala Lumpur

On-site
MYR 240,000 - 360,000
Tanzu engineer
Tanzu engineer

Encora Inc. • Kuala Lumpur

On-site
MYR 120,000 - 240,000
Tanzu engineer
Tanzu engineer

Linuxconfig • Kuala Lumpur

On-site
MYR 180,000 - 240,000
Senior DevOps Engineer
Senior DevOps Engineer

Involve Asia • Kuala Lumpur

On-site
MYR 180,000 - 300,000
OpenShift Engineer
OpenShift Engineer

Linuxconfig • Kuala Lumpur

On-site
MYR 120,000 - 160,000
VMware Cloud Foundation SRE - Platform Reliability
VMware Cloud Foundation SRE - Platform Reliability

Great Eastern • Cyberjaya

On-site
MYR 180,000 - 300,000
Senior Tanzu & Kubernetes Platform Engineer
Senior Tanzu & Kubernetes Platform Engineer

Encora Inc. • Kuala Lumpur

On-site
MYR 120,000 - 240,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Great Eastern • Kuala Lumpur

On-site
MYR 90,000 - 130,000
Senior Platform Engineer - DevOps
Senior Platform Engineer - DevOps

Endava • Kuala Lumpur

On-site
MYR 180,000 - 240,000
Share plan
Global career opportunities
Hybrid work hours
+1
Platform Architect (Senior)
Platform Architect (Senior)

INSCALE • Kuala Lumpur

On-site
MYR 180,000 - 280,000