Site Reliability Engineer

Great Eastern

Cyberjaya

On-site

MYR 180,000 - 240,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Great Eastern is seeking a Site Reliability Engineer to ensure the reliability, availability and performance of VMware Cloud Foundation (VCF) infrastructure across SG and MY. You will automate, monitor and proactively resolve issues, collaborating with Dev and Ops teams to improve platform stability and contribute to VCF roadmap.

The role requires deep VCF/NSX/vSphere expertise, hands-on scripting with Python/PowerShell, and strong automation with Ansible/Terraform, plus experience with

Qualifications

  • Bachelor’s Degree in Computer Science, IT, Computer Engineering, or related field; Master’s degree or certifications (ITIL, TOGAF, Cloud, VMware VCP) a plus.
  • Extensive knowledge of VMware Cloud Foundation (VCF) with hands-on experience of vSphere, vSAN, NSX, and vRealize Suite.
  • Minimum 15 years in IT infrastructure roles, with enterprise VMware VCF deployment experience.
  • Scripting in Python or PowerShell.
  • Hands-on automation with Ansible and Terraform.
  • Experience with monitoring/alerting: Prometheus, Grafana, Dynatrace, vRealize Operations.
  • Strong communication and cross-functional collaboration capabilities.

Responsibilities

  • Design, implement, and maintain VCF infrastructure to support GE’s requirements.
  • Manage VCF resource availability (compute/memory/storage) across clusters in SG and MY.
  • Integrate backup services using NetBackup (HotAdd, image/file backups).
  • Conduct daily health checks and monitor metrics via vROPS, vCenter, Dynatrace.
  • Analyze VCF components and perform NVA security remediation for compliance.
  • Adhere to MAS (SG) and BNC (MY) regulatory guidance; monitor threats.
  • Develop SOPs for VCF operations, recovery procedures, DR plans.
  • Apply updates/patching to ESXi hosts and vSphere components.
  • Collaborate with security to implement policies including hardening measures.
  • Drive VCF-related projects with timely delivery and stakeholder alignment.
  • Coordinate with vendors/contractors for deployment on VCF platform.

Skills

VCF expertise
NSX
vCenter
ESXi
vSphere
vRealize Suite
Python
PowerShell
Ansible
Terraform
Prometheus
Grafana
Dynatrace
Troubleshooting

Education

Bachelor’s Degree in CS/IT/CE
Master’s degree or VMware/ITIL/TOGAF/cloud cert

Tools

NSX-T
NetBackup
vROPS
Dynatrace
vRealize Operations
vRealize Automation

Job description

The Site Reliability Engineer (SRE) for VMware Cloud Foundation (VCF) focuses on ensuring the


reliability, availability, and performance of the VCF platform through automation, monitoring, and


proactive problem-solving. This role involves developing and implementing strategies to improve the platform's stability, collaborating with development and operations teams, and contributing to the overall VCF roadmap.


Key Responsibilities:


  • Design, implement, and maintain VMware Cloud Foundation (VCF) infrastructure to support GE’s organizational requirements.

  • Manage and troubleshoot VCF resource availability, including compute, memory, and storage (SAN and vSAN) up to 160 ESXi and more than 1200 VMs across multiple clusters located in both SG and MY, using tools such as NSX-T, vCenter, ESXi 8.x, and VMware vSphere Cluster availability.

  • Experience in integration with backup services using NetBackup (NBU) such as HotAdd for image backup/restore and Media to file level backup/restore to support business application VMs requirements, including full, incremental, and ad-hoc backups.

  • Conduct daily health checks and monitor VCF infrastructure metrics via vROPS, vCenter, Dynatrace to ensure optimal workload performance and timely issue resolution.

  • Analyze VCF components and perform NVA security remediation to maintain compliance across vSphere, NSX, vSAN, and other VCF elements.

  • Maintains awareness of industry trends on regulatory MAS (SG) and BNC (MY) compliance, emerging threats and technologies to understand the risk and better safeguard the company. Experience with HPSA scanning tools is a plus.

  • Develop and maintain comprehensive Standard Operating Procedures (SOPs) for VCF operations, including ESXi uptime/downtime records, VCF inventory, recovery procedures, and disaster recovery plans.

  • Apply updates, service packs and patching to ESXi hosts and vSphere components to ensure security and product currency.

  • Collaborate with the security team to implement required policies, including hardening measures to protect VCF nodes, NSX firewall, DSA on VM level etc.

  • Takes accountability in considering business and regulatory compliance risks and takes appropriate steps to mitigate the risks.

  • Execute VCF-related infrastructure projects, ensuring timely delivery and alignment with business requirements.

  • Work with vendors and third-party contractors to manage projects and implementation of VCF-related products and services. Partner with the project delivery team to identify business application requirements and support deployment on the VCF platform.


We are looking for people with


  • Bachelor’s Degree in Computer Science, Information Technology, Computer Engineering, or a related field. A Master’s degree or relevant certifications (e.g., ITIL, TOGAF, Cloud certifications or VMware VCP) is a plus.

  • Extensive Knowledge of VMware Cloud Foundation (VCF) and NSX: Strong understanding and hands-on experience with VCF components, including vSphere, vSAN, NSX, and the vRealize Suite (e.g., vRealize Automation, vRealize Operations).

  • Minimum 15 years of experience in IT infrastructure roles, with a significant portion focused on VMware VCF solutioning, hand-on deployment experiences, and be able to work on enterprise level capabilities.

  • Skilled in scripting languages such as Python or PowerShell.

  • Hand-on Experience with automation tools and frameworks, including Ansible and Terraform.

  • Familiarity with implementing monitoring and alerting solutions such as Prometheus, Grafana, Dynatrace, or vRealize Operations.

  • Demonstrated ability to troubleshoot and resolve complex technical issues effectively.

  • Strong communication skills with the ability to collaborate across various levels of stakeholders.


How you succeed


  • Champion and embody our Core Values in everyday tasks and interactions.

  • Demonstrate high level of integrity and accountability.

  • Take initiative to drive improvements and embrace change.

  • Take accountability of business and regulatory compliance risks, implementing measures to mitigate them effectively.

  • Keep abreast with industry trends, regulatory compliance, and emerging threats and technologies to understand and highlight potential concerns/ risks to safeguard our company proactively.


Who we are

Founded in 1908, Great Eastern is a well-established market leader and trusted brand in Singapore and Malaysia. With over S$100 billion in assets and more than 16 million policyholders, including 12.5 million from government schemes, it provides insurance solutions to customers through three successful distribution channels – a tied agency force, bancassurance, and financial advisory firm Great Eastern Financial Advisers. The Group also operates in Indonesia and Brunei.


The Great Eastern Life Assurance Company Limited and Great Eastern General Insurance Limited have been assigned the financial strength and counterparty credit ratings of "AA-" by S&P Global Ratings since 2010, one of the highest among Asian life insurance companies. Great Eastern's asset management subsidiary, Lion Global Investors Limited, is one of the leading asset management companies in Southeast Asia.


Great Eastern is a subsidiary of OCBC, the longest established Singapore bank, formed in 1932. It is the second largest financial services group in Southeast Asia by assets and one of the world’s most highly-rated banks, with an Aa1 rating from Moody’s and AA- by both Fitch and S&P. Recognised for its financial strength and stability, OCBC is consistently ranked among the World’s Top 50 Safest Banks by Global Finance and has been named Best Managed Bank in Singapore by The Asian Banker.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Platform Engineer
Cloud Platform Engineer

Great Eastern • Cyberjaya

On-site
MYR 180,000 - 320,000
Middleware Engineer
Middleware Engineer

Great Eastern • Kuala Lumpur

On-site
MYR 134,000 - 201,000
Senior IT Business Partner
Senior IT Business Partner

Great Eastern • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Audit Associate (Opportunity to gain experience with one of the Big 4 Accounting Firm!)
Audit Associate (Opportunity to gain experience with one of the Big 4 Accounting Firm!)

Great Eastern • Kuala Lumpur

On-site
MYR 60,000 - 90,000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Great Eastern • Kuala Lumpur

On-site
MYR 180,000 - 300,000
Intern (GELM Operations-Policy Servicing Administration)
Intern (GELM Operations-Policy Servicing Administration)

Great Eastern • Kuala Lumpur

On-site
MYR 60,000 - 90,000
Senior Full Stack Software Engineer
Senior Full Stack Software Engineer

Great Eastern • Kuala Lumpur

On-site
MYR 180,000 - 280,000
Health insurance
Performance bonuses
Senior Infrastructure Project Manager
Senior Infrastructure Project Manager

Great Eastern • Kuala Lumpur

On-site
MYR 240,000 - 360,000
Manager, Financial Risk
Manager, Financial Risk

Great Eastern • Kuala Lumpur

On-site
MYR 180,000 - 260,000
Lead Database Platform Engineer
Lead Database Platform Engineer

Great Eastern • Cyberjaya

On-site
MYR 180,000 - 260,000