Get more replies from employers
Send a job-specific resume in minutes.
Great Eastern is seeking a Site Reliability Engineer to ensure the reliability and performance of VMware Cloud Foundation infrastructure. The role emphasizes automation, monitoring, and proactive problem-solving to support enterprise workloads.
You will design, implement, and maintain VCF infrastructure, manage resources across clusters and sites, and collaborate with security and operations teams to sustain service levels.
Jora Malaysia will close on 9th September 2026. Thank you for being with us, we are cheering you on as you continue your career journey.
A Site Reliability Engineer (SRE) for VMware Cloud Foundation (VCF) focuses on ensuring the reliability, availability, and performance of the VCF platform through automation, monitoring, and proactive problem-solving. This role involves developing and implementing strategies to improve the platform's stability, collaborating with development and operations teams, and contributing to the overall VCF roadmap.
Experience to design, implement, and maintain VMware Cloud Foundation (VCF) infrastructure to support GE’s organizational requirements.
Manage and troubleshoot VCF resource availability, including compute, memory, and storage (SAN and vSAN) up to 160 ESXi and more than 1200 VMs across multiple clusters located in both SG and MY, using tools such as NSX-T, vCenter, ESXi 8.x, and VMware vSphere Cluster availability.
Experience in integration with backup services using NetBackup (NBU) such as HotAdd for image backup/restore and Media to file level backup/restore to support business application VMs requirements, including full, incremental, and ad-hoc backups.
Conduct daily health checks and monitor VCF infrastructure metrics via vROPS, vCenter, Dynatrace to ensure optimal workload performance and timely issue resolution.
Analyze VCF components and perform NVA security remediation to maintain compliance across vSphere, NSX, vSAN, and other VCF elements. Maintains awareness of industry trends on regulatory MAS (SG) and BNC (MY) compliance, emerging threats and technologies to understand the risk and better safeguard the company. Experience with HPSA scanning tools is a plus.
Develop and maintain comprehensive Standard Operating Procedures (SOPs) for VCF operations, including ESXi uptime/downtime records, VCF inventory, recovery procedures, and disaster recovery plans.
Apply updates, service packs and patching to ESXi hosts and vSphere components to ensure security and product currency.
Collaborate with the security team to implement required policies, including hardening measures to protect VCF nodes, NSX firewall, DSA on VM level etc. Takes accountability in considering business and regulatory compliance risks and takes appropriate steps to mitigate the risks.
Execute VCF-related infrastructure projects, ensuring timely delivery and alignment with business requirements.
Work with vendors and third-party contractors to manage projects and implementation of VCF-related products and services. Partner with the project delivery team to identify business application requirements and support deployment on the VCF platform.
Strong understanding and hands-on experience with VCF components, including vSphere, vSAN, NSX, and the vRealize Suite (e.g., vRealize Automation, vRealize Operations).
15 years of experience in IT infrastructure roles, with a significant portion focused on VMware VCF solutioning, hand-on deployment experiences, and be able to work on enterprise level capabilities.
Skilled in scripting languages such as Python or PowerShell.
Hand-on Experience with automation tools and frameworks, including Ansible and Terraform.
Familiarity with implementing monitoring and alerting solutions such as Prometheus, Grafana, Dynatrace, or vRealize Operations.
Demonstrated ability to troubleshoot and resolve complex technical issues effectively.
Strong communication skills with the ability to collaborate across various levels of stakeholders.