Turn this role into an interview — a resume and cover letter built around what this employer wants.
Raydian Cloud Sdn Bhd in Malaysia seeks a junior IT support engineer to provide first-line operational support for GPU compute infrastructure in a customer data centre. You will monitor GPU servers, check health, assist remote engineers and escalate issues as needed.
The role suits fresh graduates with a foundation in Linux and server hardware who want practical experience supporting high-density GPU platforms in a managed services environment. English communication is required.
This role provides first-line operational support for GPU compute infrastructure in a customer data centre. The engineer monitors GPU servers and supporting systems, performs approved routine tasks, completes initial diagnosis, and escalates faults that require advanced troubleshooting, privileged access or vendor intervention. The position suits a fresh graduate or junior engineer with a foundation in Linux and server hardware who wants practical experience supporting high-density GPU platforms in a managed services environment.
Monitor GPU compute nodes, management servers, operating systems, storage connections, backup jobs and related infrastructure using the customer monitoring platforms.
Respond to system, hardware, GPU, power and environmental alerts within the agreed service levels and operational procedures.
Perform first-level checks through the server management interface, such as BMC, iDRAC, iLO or IPMI, to confirm power state, temperature, fan, power supply, memory, disk and PCIe health.
Check GPU visibility and health using approved tools, including GPU utilisation, temperature, power, memory, ECC and Xid alerts where the platform exposes this information.
Collect diagnostic evidence such as nvidia-smi output, system logs, hardware event logs, driver and CUDA versions, screenshots and timestamps for L2/L3 or vendor review.
Perform approved routine actions such as node health checks, service restarts, power cycles, log collection, account administration, storage mount checks and backup verification.
Check cluster node status in approved management platforms, such as Slurm or Kubernetes, and follow the documented procedure for draining, restarting or returning a node to service.
Provide remote hands support by checking indicators, tracing cables, confirming asset labels and assisting remote engineers during diagnosis and maintenance.
Support rack and stack, server installation, cabling and approved component replacement under a work order and supervision.
Record incidents, requests, diagnostic results, actions and outcomes accurately in the ticketing system and shift handover.
Diploma or bachelor degree in Information Technology, Computer Science, Computer Engineering or a related field.
Fresh graduates and candidates with up to two years of IT support, infrastructure support or data centre experience are welcome.
Basic Linux command-line and operating system troubleshooting skills. Windows Server knowledge is an advantage.
Basic understanding of server components, including CPU, memory, storage, network adapters, power supplies and PCIe devices.
Basic awareness of GPU servers and the relationship between the operating system, GPU driver and compute workload.
Ability to use system logs and standard diagnostic tools to investigate common hardware, operating system and connectivity issues.
Ability to follow procedures and checklists carefully in a controlled production environment.
Working proficiency in English. Bahasa Malaysia or Mandarin is an advantage for customer communication.
Clear communication, teamwork and customer service skills.
Legal eligibility to work in Malaysia.