Senior Data Centre Operations Engineer

Oxydata Software Sdn Bhd

Malaysia

On-site

MYR 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Oxydata Software Sdn Bhd in Malaysia seeks a Senior Data Centre Operations Engineer to manage and support server and GPU infrastructure in large-scale environments, ensuring reliable AI and data centre operations.

The role requires strong Linux administration, hardware diagnostics, and hands-on experience with GPU clusters, RAID, iDRAC/ILO, and vendor coordination to maintain uptime and performance across a state-of-the-art facility in Johor Bahru.

Qualifications

  • Bachelor's degree in Computer Science, Information Systems, Electronic Engineering or a related field.
  • Willingness to travel and support overtime, night shifts, on-call duties or weekend work when required.
  • At least 5-7 years of experience in server operations.
  • Experience using NVIDIA diagnostic tools, including NVIDIA-SMI and DCGM.
  • Strong Linux administration and troubleshooting knowledge.
  • Good hardware troubleshooting and problem-solving skills.
  • Ability to work effectively with internal teams and external vendors.
  • Good technical communication skills in English.

Responsibilities

  • Set up, configure, and troubleshoot server hardware, including CPUs, memory, storage, RAID, NICs, and power supplies.
  • Monitor server health and review IPMI, BMC, and operating system logs.
  • Manage BIOS, BMC, iDRAC, iLO, and firmware upgrades.
  • Manage RAID configurations and monitor SSD and NVMe health.
  • Troubleshoot GPU servers and replace faulty hardware components.
  • Support GPU cluster performance, network topology, and system stability.
  • Collaborate with networking, storage and virtualization teams to resolve issues.
  • Automate firmware upgrades, inspections and hardware alert management.
  • Prepare technical guides, troubleshooting documentation and SOPs.
  • Coordinate with hardware vendors and manage RMAs and spare parts.
  • Monitor rack power, temperature, and air-cooling conditions.

Skills

Linux administration
Hardware troubleshooting
NVIDIA diagnostics (NVIDIA-SMI/DCGM)
Server operations
Technical communication in English
Automation (Shell/Ansible)

Education

Bachelor's degree in Computer Science, Information Systems, Electronic Engineering or related field

Tools

NVIDIA diagnostic tools (NVIDIA-SMI/DCGM)

Job description

Senior Data Centre Operations Engineer

Location: Senai, Johor, Malaysia
Work Mode: Onsite

Employment type: Permanent

Our client is a leading specialist in the repair and maintenance of high-end AI computing infrastructure, with a state-of-the-art facility located in Johor Bahru, Malaysia. They are dedicated to providing mission-critical support and have established a reputation for precision and reliability in the Southeast Asian market. With a strong focus on transparency and accountability, they ensure that every repair process is documented and approved by clients, maintaining a high standard of service excellence.

We are seeking an experienced Senior Data Centre Operations Engineer to manage and support server and GPU infrastructure in large-scale environments, ensuring reliable AI and data centre operations.

Responsibilities
  • Set up, configure, and troubleshoot server hardware, including CPUs, memory, storage, RAID, NICs, and power supplies.
  • Monitor server health and review IPMI, BMC, and operating system logs.
  • Manage BIOS, BMC, iDRAC, iLO, and firmware upgrades.
  • Manage RAID configurations and monitor SSD and NVMe health.
  • Troubleshoot GPU servers and replace faulty hardware components.
  • Support GPU cluster performance, network topology, and system stability.
  • Collaborate with networking, storage, and virtualisation teams to resolve technical issues.
  • Automate routine tasks such as firmware upgrades, inspections, and hardware alert management.
  • Prepare technical guides, troubleshooting documentation, and standard operating procedures.
  • Coordinate with hardware vendors and manage RMAs and spare parts.
  • Monitor rack power, temperature, and air-cooling conditions.
Requirements
Must-have:
  • Bachelor's degree in Computer Science, Information Systems, Electronic Engineering or a related field.
  • Willingness to travel and support overtime, night shifts, on-call duties or weekend work when required.
  • At least 5-7 years of experience in server operations.
  • Experience using NVIDIA diagnostic tools, including NVIDIA-SMI and DCGM.
  • Strong Linux administration and troubleshooting knowledge.
  • Good hardware troubleshooting and problem-solving skills.
  • Ability to work effectively with internal teams and external vendors.
  • Good technical communication skills in English.
Nice-to-have:
  • Experience operating large-scale GPU clusters.
  • Experience with Shell or Ansible automation.
  • Familiarity with InfiniBand, RoCE, RDMA networking and optical modules.
  • Exposure to air-cooled server environments.
  • Experience with liquid-cooled servers.
  • Knowledge of Ceph, KVM, VMware or hyper-converged infrastructure.
  • Relevant certifications such as RHCE, RHCA, CompTIA Server+ or server vendor certifications.
Education:
  • Bachelor's degree in Computer Science, Information Systems, Electronic Engineering or a related field.
Why Join Us
  • Be part of a dynamic team at the forefront of data centre technology, supporting mission-critical AI infrastructure for leading enterprises.
  • Opportunity to work with advanced hardware, collaborate with skilled professionals, and contribute to the reliability of high-performance computing environments across the region.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Hardware Field Service Engineer
GPU Hardware Field Service Engineer

Oxydata Software Sdn Bhd • Kulai

On-site
MYR 60,000 - 120,000
System Engineer – Infrastructure (AI & HPC Systems)
System Engineer – Infrastructure (AI & HPC Systems)

Neuron Solutions Sdn. Bhd. • Johor Bahru

On-site
MYR 90,000 - 150,000
Monetary compensation
Senior Data Center Ops Engineer — GPU & AI Infra
Senior Data Center Ops Engineer — GPU & AI Infra

Oxydata Software Sdn Bhd • Malaysia

On-site
MYR 120,000 - 180,000
GPU Hardware Field Service Engineer
GPU Hardware Field Service Engineer

Oxydata Software Sdn Bhd • Malaysia

On-site
MYR 90,000 - 150,000
Senior AI Data Centre Network Engineer
Senior AI Data Centre Network Engineer

YTL AI Cloud • Kulai

On-site
MYR 180,000 - 280,000
Senior AI Network & Security Engineer (Johor Bahru)
Senior AI Network & Security Engineer (Johor Bahru)

Techstreet • Johor Bahru

On-site
MYR 180,000 - 300,000
Senior Data Center Operations Engineer, Infrastructure (Kulai, Johor)
Senior Data Center Operations Engineer, Infrastructure (Kulai, Johor)

Shopee • Kuala Lumpur

On-site
MYR 180,000 - 240,000
Technical Manager - GPU Cloud & AI Infrastructure
Technical Manager - GPU Cloud & AI Infrastructure

Risewave Consulting, Inc. • Kuala Lumpur

On-site
MYR 180,000 - 280,000
Data Centre Operations Manager
Data Centre Operations Manager

Randstad Malaysia • Kulai

On-site
Medical insurance and hospitalization coverage
Attractive performance bonus structure
Senior AI Network - Security Engineer
Senior AI Network - Security Engineer

Techstreet Malaysia • Johor Bahru

On-site
MYR 180,000 - 280,000