System Engineer

Raydian Cloud Sdn Bhd

Kuala Lumpur

On-site

MYR 18,000 - 30,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Raydian Cloud Sdn Bhd in Malaysia seeks a junior IT support engineer to provide first-line operational support for GPU compute infrastructure in a customer data centre. You will monitor GPU servers, check health, assist remote engineers and escalate issues as needed.

The role suits fresh graduates with a foundation in Linux and server hardware who want practical experience supporting high-density GPU platforms in a managed services environment. English communication is required.

Qualifications

  • Fresh graduates or up to 2 years within IT/support/data centre environments.
  • Basic Linux command-line and OS troubleshooting.
  • Windows Server knowledge is a plus.
  • Solid understanding of server components (CPU, memory, storage, NICs).
  • Awareness of GPU servers and OS-driver-workload relationships.
  • Ability to use logs and diag tools for hardware/OS/connectivity issues.
  • Ability to follow procedures in a controlled production environment.
  • English communication required; Malay/ Mandarin an advantage.

Responsibilities

  • Monitor GPU compute nodes and related infrastructure using monitoring platforms.
  • Respond to system, hardware, power, and environmental alerts within SLA.
  • Perform first-level checks through BMC/iDRAC/iLO/IPMI to confirm health.
  • Check GPU visibility and health using approved tools and metrics.
  • Collect diagnostic evidence for L2/L3 or vendor review.
  • Record incidents and actions accurately in the ticketing system.
  • Provide remote hands support and assist remote engineers during diagnosis.
  • Support rack and stack, cabling, and component replacement under work orders.

Skills

Linux basics
Troubleshooting
Server hardware knowledge
English communication
Attention to detail
Teamwork

Education

Diploma or bachelor’s degree in IT/CS/Engineering

Tools

BMC/iDRAC/iLO

Job description

This role provides first-line operational support for GPU compute infrastructure in a customer data centre. The engineer monitors GPU servers and supporting systems, performs approved routine tasks, completes initial diagnosis, and escalates faults that require advanced troubleshooting, privileged access or vendor intervention. The position suits a fresh graduate or junior engineer with a foundation in Linux and server hardware who wants practical experience supporting high-density GPU platforms in a managed services environment.

Key responsibilities

Monitor GPU compute nodes, management servers, operating systems, storage connections, backup jobs and related infrastructure using the customer monitoring platforms.

Respond to system, hardware, GPU, power and environmental alerts within the agreed service levels and operational procedures.

Perform first-level checks through the server management interface, such as BMC, iDRAC, iLO or IPMI, to confirm power state, temperature, fan, power supply, memory, disk and PCIe health.

Check GPU visibility and health using approved tools, including GPU utilisation, temperature, power, memory, ECC and Xid alerts where the platform exposes this information.

Collect diagnostic evidence such as nvidia-smi output, system logs, hardware event logs, driver and CUDA versions, screenshots and timestamps for L2/L3 or vendor review.

Perform approved routine actions such as node health checks, service restarts, power cycles, log collection, account administration, storage mount checks and backup verification.

Check cluster node status in approved management platforms, such as Slurm or Kubernetes, and follow the documented procedure for draining, restarting or returning a node to service.

Provide remote hands support by checking indicators, tracing cables, confirming asset labels and assisting remote engineers during diagnosis and maintenance.

Support rack and stack, server installation, cabling and approved component replacement under a work order and supervision.

Record incidents, requests, diagnostic results, actions and outcomes accurately in the ticketing system and shift handover.

About you

Diploma or bachelor degree in Information Technology, Computer Science, Computer Engineering or a related field.

Fresh graduates and candidates with up to two years of IT support, infrastructure support or data centre experience are welcome.

Basic Linux command-line and operating system troubleshooting skills. Windows Server knowledge is an advantage.

Basic understanding of server components, including CPU, memory, storage, network adapters, power supplies and PCIe devices.

Basic awareness of GPU servers and the relationship between the operating system, GPU driver and compute workload.

Ability to use system logs and standard diagnostic tools to investigate common hardware, operating system and connectivity issues.

Ability to follow procedures and checklists carefully in a controlled production environment.

Working proficiency in English. Bahasa Malaysia or Mandarin is an advantage for customer communication.

Clear communication, teamwork and customer service skills.

Legal eligibility to work in Malaysia.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network Engineer
Network Engineer

Raydian Cloud Sdn Bhd • Kuala Lumpur

On-site
MYR 3,200 - 6,400
GPU Hardware Field Service Engineer
GPU Hardware Field Service Engineer

Oxydata Software Sdn Bhd • Malaysia

On-site
MYR 90,000 - 150,000
MA ENGINEER
MA ENGINEER

QUETTA BYTE SDN. BHD. • Johor Bahru

On-site
MYR 60,000 - 90,000
Senior Data Centre Operations Engineer
Senior Data Centre Operations Engineer

Oxydata Software Sdn Bhd • Malaysia

On-site
MYR 120,000 - 180,000
Data Centre GPU Infrastructure Engineer (Based KL)
Data Centre GPU Infrastructure Engineer (Based KL)

CloudEngine Digital Co., Ltd • Kuala Lumpur

On-site
MYR 67,000 - 112,000
System Engineer – Infrastructure (AI & HPC Systems)
System Engineer – Infrastructure (AI & HPC Systems)

Neuron Solutions Sdn. Bhd. • Johor Bahru

On-site
MYR 90,000 - 150,000
Monetary compensation
Technical Manager - GPU Cloud & AI Infrastructure
Technical Manager - GPU Cloud & AI Infrastructure

Risewave Consulting, Inc. • Kuala Lumpur

On-site
MYR 180,000 - 280,000
GPU Data Center Infrastructure Engineer — Kuala Lumpur
GPU Data Center Infrastructure Engineer — Kuala Lumpur

CloudEngine Digital Co., Ltd • Kuala Lumpur

On-site
MYR 67,000 - 112,000
Cloud GPU Operations Engineer (L1/L2) – 24/7 NOC & Support
Cloud GPU Operations Engineer (L1/L2) – 24/7 NOC & Support

Straitdeer Pte. Ltd. • Cyberjaya

On-site
MYR 70,000 - 100,000
Cloud Operations & Support Engineer (L1 / L2)
Cloud Operations & Support Engineer (L1 / L2)

Straitdeer Pte. Ltd. • Cyberjaya

On-site
MYR 70,000 - 100,000