Senior Network Engineer

Raydian Cloud

Kuala Lumpur

On-site

MYR 120,000 - 180,000

Full time

43 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Raydian Cloud in Kuala Lumpur is seeking an AI Cluster Network Engineer to design, deploy, and optimize high-performance networks for AI/GPU computing and GPUaaS environments. You will work with NVIDIA Spectrum-X, InfiniBand, RoCEv2, and related technologies to deliver multi-tenant, secure, and scalable cluster networks.

The role focuses on RDMA tuning, NCCL traffic patterns, automation, observability, and production support across the Asia-Pacific region.

Qualifications

  • Bachelor-level degree in CS/IT/Electronics/Telecommunications or related field.
  • Minimum 5 years in large-scale data center networks, deployment and troubleshooting.
  • Minimum 3 years with InfiniBand, RoCEv2, and RDMA technologies.
  • Experience supporting AI/GPU clusters, HPC, or large distributed systems.

Responsibilities

  • Design and deploy AI cluster networks using NVIDIA Spectrum-X and InfiniBand.
  • Build leaf-spine, multi-tenant GPUaaS networks with proper segmentation.
  • Tune and optimize RDMA, NCCL traffic, congestion control and QoS.
  • Troubleshoot switches, NICs, DPUs, and Linux networking stack.
  • Collaborate with compute/storage teams; perform performance validation.
  • Develop automation, monitoring, runbooks, and incident response workflows.
  • Provide POCs, benchmarks, and support production deployments in APAC.

Skills

BGP/ECMP/VXLAN
QoS/high availability
RDMA/NICs/RoCEv2
AI/GPU cluster networking
NCCL & HPC networking
Networking troubleshooting

Education

Bachelor’s degree in CS/IT/Electronics/Telecommunications

Tools

NVIDIA Spectrum-X
NVIDIA ConnectX NICs
NVIDIA BlueField DPUs
InfiniBand hardware
Arista/Cisco/Juniper gear
tcpdump/ethtool

Job description

1. Key Responsibilities
1.1 AI Cluster Network Architecture Design & Deployment
  • Design and deploy AI cluster network architectures based onNVIDIA Spectrum-X EthernetandNVIDIA Quantum InfiniBandtechnologies.
  • BuildLeaf-Spine and Clos network architecturescovering GPU compute networks, storage networks, management networks, and cross-site connectivity.
  • Develop standard reference architectures forsingle-tenant and multi-tenant GPUaaS environments, including network segmentation, routing domains, and business zones.
  • Evaluate and select network hardware based on performance, cost, reliability, scalability, and security requirements.
1.2 RDMA, RoCEv2 & InfiniBand Operations and Optimization
  • Configure, operate, troubleshoot, and optimizeRoCEv2 and InfiniBand lossless networks.
  • Perform in-depth RDMA performance tuning for distributed AI/ML workloads, includingNCCL, MPI, and inter-GPU traffic.
  • Manage and optimizecongestion control, ECN, PFC, QoS, and adaptive routing.
  • Troubleshoot communication bottlenecks caused by switches, NICs, DPUs, optical transceivers, firmware, and the Linux networking stack.
1.3 Cluster Integration & Performance Validation
  • Work closely with compute, storage, and data center teams to integrate and commission GPUs, NICs, DPUs, storage nodes, and related infrastructure.
  • Conduct network and cluster performance testing using tools such asib_write_bw, iperf, NCCL test suites, and other benchmarking tools.
  • Establish performance baselines covering latency, bandwidth, packet loss, CRC errors, and other critical network metrics.
  • Support customerPOCs, benchmark testing, technical validation, and production deployment.
1.4 Multi-Tenant Network Services
  • Implement tenant isolation usingVLAN, VRF, VXLAN/EVPN, and related data center networking technologies.
  • Design and deploy customer connectivity solutions including dedicated circuits, cloud interconnects,IPsec/VPN, and other secure access services.
  • Operate network services including firewalls, load balancers, DNS, DHCP, and out-of-band management.
  • Develop standardized deployment and onboarding templates to improve tenant provisioning efficiency.
1.5 Automation, Monitoring & Operations
  • Develop network automation frameworks usingPython/Go, Ansible, CI/CD, and Infrastructure as Code (IaC).
  • ImplementZero-Touch Provisioning (ZTP), configuration validation, automated deployment, and version rollback capabilities.
  • Build comprehensive telemetry and monitoring covering port status, optical power levels, congestion, RDMA metrics, and other critical infrastructure indicators.
  • Develop and maintain operational runbooks, escalation procedures, incident response processes, andRoot Cause Analysis (RCA)templates.
  • Support24×7 production operationsand participate in on-call and emergency incident response.
1.6 Project Delivery & Technical Documentation
  • Participate in data center planning, including rack layout, network cabling, IP addressing, hardware BOMs, and infrastructure design.
  • ProduceHigh-Level Design (HLD), Low-Level Design (LLD), Method of Procedure (MOP), test plans, test reports, and other technical documentation.
  • Provide technical support for major production incidents and participate in project deployments across theAsia-Pacific region.
  • Coordinate with data center operators, hardware vendors, customers, and internal technical teams to ensure successful project delivery.
2. Requirements
2.1 Essential Requirements
  • Bachelor’s degree or above inComputer Science, Information Technology, Electronics, Telecommunications, or a related discipline.
  • Minimum5 years of experiencein large-scale data center network design, deployment, and troubleshooting.
  • Minimum3 years of hands-on experiencewithInfiniBand, RoCEv2, and RDMA.
  • Proven experience supportingAI/GPU clusters, High-Performance Computing (HPC), or large-scale distributed computing environments.
2.2 Required Technical Skills
  • Strong knowledge ofBGP, ECMP, VXLAN/EVPN, QoS, high availability, and modern data center networking technologies.
  • Hands-on experience withNVIDIA/Mellanoxnetworking products, including:
  • NVIDIA Spectrum switches
  • NVIDIA ConnectX NICs
  • NVIDIA BlueField DPUs
  • Familiarity with networking equipment from major vendors such asArista, Cisco, and Juniper.
  • Strong Linux networking troubleshooting skills using tools such astcpdump, ethtool, and related utilities.
  • Experience with network monitoring and observability platforms such asPrometheus, Grafana, and ELK.
2.3 Preferred / Additional Qualifications
  • Hands-on experience deployingNVIDIA DGX, HGX, GB-series AI systems, Spectrum-X, or Quantum InfiniBand.
  • Knowledge ofNCCL, GPUDirect RDMA, and large-scale AI/ML workload traffic patterns, including MoE workloads.
  • Familiarity withKubernetes networking, Slurm, parallel file systems, NVMe-oF, and related AI/HPC technologies.
  • Experience withPython/Go, Ansible, CI/CD, and network automation.
  • NVIDIA, Arista, or other relevant vendor certifications are an advantage.
3. Performance Indicators & Additional Requirements
3.1 Key Performance Indicators
  • Deliver network projects on schedule with complete and accurate cabling, configuration, testing, and acceptance documentation.
  • EnsureRDMA and NCCL performancemeets defined customer and production requirements.
  • Respond rapidly to network incidents and accurately identify root causes, with complete RCA documentation and corrective action plans.
  • Continuously improve network automation, operational efficiency, tenant security, and network isolation.
3.2 Additional Requirements
  • Willingness to travel frequently acrossAsia-Pacificcountries.
  • Able to support scheduled night-time maintenance, emergency troubleshooting, and on-call operations when required.
  • Comfortable working in a fast-paced project environment with multiple stakeholders.
  • Strong communication and coordination skills when working with data center operators, hardware vendors, enterprise customers, and internal engineering teams.
4. Job Summary

TheAI Cluster Network Engineeris a key technical role responsible for designing, deploying, optimizing, and operating high-performance networks forAI/GPU computing and GPUaaS environments.

The role is deeply focused on NVIDIA’s advanced networking technologies, includingSpectrum-X Ethernet, Quantum InfiniBand, RoCEv2, RDMA, and high-performance cluster networking. The successful candidate will combine strong traditional data center networking expertise with hands-on experience in AI/HPC networking and distributed workload optimization.

In addition to network architecture and performance tuning, the role coversmulti-tenant network isolation, automation, observability, production operations, customer POCs, and project deliveryacross the Asia-Pacific region.

This position is ideal for a senior network engineer who wants to specialize inAI infrastructure, GPU clusters, high-performance computing, and next-generation data center networking, with a strong focus on commercial, production-grade GPUaaS deployments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Network & Security Engineer - Data Centre
AI Network & Security Engineer - Data Centre

Neuron Solutions Sdn. Bhd. • Johor

On-site
MYR 120,000 - 180,000
Senior AI Network - Security Engineer
Senior AI Network - Security Engineer

Techstreet Malaysia • Johor Bahru

On-site
MYR 180,000 - 280,000
Senior AI Network & Security Engineer (Johor Bahru)
Senior AI Network & Security Engineer (Johor Bahru)

Techstreet • Johor Bahru

On-site
MYR 180,000 - 300,000
Senior AI Data Centre Network Engineer
Senior AI Data Centre Network Engineer

YTL AI Cloud • Kulai

On-site
MYR 180,000 - 280,000
Technical Manager - GPU Cloud & AI Infrastructure
Technical Manager - GPU Cloud & AI Infrastructure

Risewave Consulting, Inc. • Kuala Lumpur

On-site
MYR 180,000 - 280,000
AI Cluster Network Engineer for GPU-Accelerated HPC
AI Cluster Network Engineer for GPU-Accelerated HPC

Raydian Cloud • Kuala Lumpur

On-site
MYR 120,000 - 180,000
AI Engineer SNS Network Right Choice with the Right People
AI Engineer SNS Network Right Choice with the Right People

SNS Network (M) Sdn. Bhd. • Petaling Jaya

On-site
MYR 180,000 - 260,000
AI Infrastructure & Orchestration Lead
AI Infrastructure & Orchestration Lead

SNS Network (M) Sdn. Bhd. • Petaling Jaya

On-site
MYR 180,000 - 260,000
AI Cloud Network Architect
AI Cloud Network Architect

Bitdeer Technologies Group • Cyberjaya

Hybrid
MYR 240,000 - 360,000
AI Data Centre Network & Security Engineer
AI Data Centre Network & Security Engineer

Neuron Solutions Sdn. Bhd. • Johor

On-site
MYR 120,000 - 180,000