AI Cluster Network Engineer for GPU-Accelerated HPC

Raydian Cloud

Kuala Lumpur

On-site

MYR 120,000 - 180,000

Full time

22 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Raydian Cloud in Kuala Lumpur is seeking an AI Cluster Network Engineer to design, deploy, and optimize high-performance networks for AI/GPU computing and GPUaaS environments. You will work with NVIDIA Spectrum-X, InfiniBand, RoCEv2, and related technologies to deliver multi-tenant, secure, and scalable cluster networks.

The role focuses on RDMA tuning, NCCL traffic patterns, automation, observability, and production support across the Asia-Pacific region.

Qualifications

  • Bachelor-level degree in CS/IT/Electronics/Telecommunications or related field.
  • Minimum 5 years in large-scale data center networks, deployment and troubleshooting.
  • Minimum 3 years with InfiniBand, RoCEv2, and RDMA technologies.
  • Experience supporting AI/GPU clusters, HPC, or large distributed systems.

Responsibilities

  • Design and deploy AI cluster networks using NVIDIA Spectrum-X and InfiniBand.
  • Build leaf-spine, multi-tenant GPUaaS networks with proper segmentation.
  • Tune and optimize RDMA, NCCL traffic, congestion control and QoS.
  • Troubleshoot switches, NICs, DPUs, and Linux networking stack.
  • Collaborate with compute/storage teams; perform performance validation.
  • Develop automation, monitoring, runbooks, and incident response workflows.
  • Provide POCs, benchmarks, and support production deployments in APAC.

Skills

BGP/ECMP/VXLAN
QoS/high availability
RDMA/NICs/RoCEv2
AI/GPU cluster networking
NCCL & HPC networking
Networking troubleshooting

Education

Bachelor’s degree in CS/IT/Electronics/Telecommunications

Tools

NVIDIA Spectrum-X
NVIDIA ConnectX NICs
NVIDIA BlueField DPUs
InfiniBand hardware
Arista/Cisco/Juniper gear
tcpdump/ethtool

Job description

Raydian Cloud in Kuala Lumpur is seeking an AI Cluster Network Engineer to design, deploy, and optimize high-performance networks for AI/GPU computing and GPUaaS environments. You will work with NVIDIA Spectrum-X, InfiniBand, RoCEv2, and related technologies to deliver multi-tenant, secure, and scalable cluster networks.

The role focuses on RDMA tuning, NCCL traffic patterns, automation, observability, and production support across the Asia-Pacific region.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI GPU Data Centre Network Engineer
Senior AI GPU Data Centre Network Engineer

YTL AI Cloud • Kulai

On-site
MYR 180,000 - 280,000
Senior Network Engineer
Senior Network Engineer

Raydian Cloud • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Senior AI Data Centre Network Engineer
Senior AI Data Centre Network Engineer

YTL AI Cloud • Kulai

On-site
MYR 180,000 - 280,000
Technical Manager - GPU Cloud & AI Infrastructure
Technical Manager - GPU Cloud & AI Infrastructure

Risewave Consulting, Inc. • Kuala Lumpur

On-site
MYR 180,000 - 280,000
AI Networking & Security Lead for Data Centers
AI Networking & Security Lead for Data Centers

Techstreet Malaysia • Johor Bahru

On-site
MYR 180,000 - 280,000
Infrastructure Systems Engineer for AI & HPC Clusters
Infrastructure Systems Engineer for AI & HPC Clusters

Neuron Solutions Sdn. Bhd. • Johor Bahru

On-site
MYR 90,000 - 150,000
Monetary compensation
Senior AI Cloud Network Architect — Hyperscale GPU Clusters
Senior AI Cloud Network Architect — Hyperscale GPU Clusters

Bitdeer (NASDAQ: BTDR) • Cyberjaya

On-site
MYR 180,000 - 320,000
AI Infra Engineer: GPU Cloud & ML Platform
AI Infra Engineer: GPU Cloud & ML Platform

Tencent • Kuala Lumpur

On-site
MYR 120,000 - 180,000
Senior AI Network - Security Engineer
Senior AI Network - Security Engineer

Techstreet Malaysia • Johor Bahru

On-site
MYR 180,000 - 280,000
AI/HPC Data Center Infrastructure Engineer
AI/HPC Data Center Infrastructure Engineer

Bitdeer • Johor Bahru

On-site
MYR 89,000 - 156,000