1. Key Responsibilities
1.1 AI Cluster Network Architecture Design & Deployment
- Design and deploy AI cluster network architectures based onNVIDIA Spectrum-X EthernetandNVIDIA Quantum InfiniBandtechnologies.
- BuildLeaf-Spine and Clos network architecturescovering GPU compute networks, storage networks, management networks, and cross-site connectivity.
- Develop standard reference architectures forsingle-tenant and multi-tenant GPUaaS environments, including network segmentation, routing domains, and business zones.
- Evaluate and select network hardware based on performance, cost, reliability, scalability, and security requirements.
1.2 RDMA, RoCEv2 & InfiniBand Operations and Optimization
- Configure, operate, troubleshoot, and optimizeRoCEv2 and InfiniBand lossless networks.
- Perform in-depth RDMA performance tuning for distributed AI/ML workloads, includingNCCL, MPI, and inter-GPU traffic.
- Manage and optimizecongestion control, ECN, PFC, QoS, and adaptive routing.
- Troubleshoot communication bottlenecks caused by switches, NICs, DPUs, optical transceivers, firmware, and the Linux networking stack.
1.3 Cluster Integration & Performance Validation
- Work closely with compute, storage, and data center teams to integrate and commission GPUs, NICs, DPUs, storage nodes, and related infrastructure.
- Conduct network and cluster performance testing using tools such asib_write_bw, iperf, NCCL test suites, and other benchmarking tools.
- Establish performance baselines covering latency, bandwidth, packet loss, CRC errors, and other critical network metrics.
- Support customerPOCs, benchmark testing, technical validation, and production deployment.
1.4 Multi-Tenant Network Services
- Implement tenant isolation usingVLAN, VRF, VXLAN/EVPN, and related data center networking technologies.
- Design and deploy customer connectivity solutions including dedicated circuits, cloud interconnects,IPsec/VPN, and other secure access services.
- Operate network services including firewalls, load balancers, DNS, DHCP, and out-of-band management.
- Develop standardized deployment and onboarding templates to improve tenant provisioning efficiency.
1.5 Automation, Monitoring & Operations
- Develop network automation frameworks usingPython/Go, Ansible, CI/CD, and Infrastructure as Code (IaC).
- ImplementZero-Touch Provisioning (ZTP), configuration validation, automated deployment, and version rollback capabilities.
- Build comprehensive telemetry and monitoring covering port status, optical power levels, congestion, RDMA metrics, and other critical infrastructure indicators.
- Develop and maintain operational runbooks, escalation procedures, incident response processes, andRoot Cause Analysis (RCA)templates.
- Support24×7 production operationsand participate in on-call and emergency incident response.
1.6 Project Delivery & Technical Documentation
- Participate in data center planning, including rack layout, network cabling, IP addressing, hardware BOMs, and infrastructure design.
- ProduceHigh-Level Design (HLD), Low-Level Design (LLD), Method of Procedure (MOP), test plans, test reports, and other technical documentation.
- Provide technical support for major production incidents and participate in project deployments across theAsia-Pacific region.
- Coordinate with data center operators, hardware vendors, customers, and internal technical teams to ensure successful project delivery.
2. Requirements
2.1 Essential Requirements
- Bachelor’s degree or above inComputer Science, Information Technology, Electronics, Telecommunications, or a related discipline.
- Minimum5 years of experiencein large-scale data center network design, deployment, and troubleshooting.
- Minimum3 years of hands-on experiencewithInfiniBand, RoCEv2, and RDMA.
- Proven experience supportingAI/GPU clusters, High-Performance Computing (HPC), or large-scale distributed computing environments.
2.2 Required Technical Skills
- Strong knowledge ofBGP, ECMP, VXLAN/EVPN, QoS, high availability, and modern data center networking technologies.
- Hands-on experience withNVIDIA/Mellanoxnetworking products, including:
- NVIDIA Spectrum switches
- NVIDIA ConnectX NICs
- NVIDIA BlueField DPUs
- Familiarity with networking equipment from major vendors such asArista, Cisco, and Juniper.
- Strong Linux networking troubleshooting skills using tools such astcpdump, ethtool, and related utilities.
- Experience with network monitoring and observability platforms such asPrometheus, Grafana, and ELK.
2.3 Preferred / Additional Qualifications
- Hands-on experience deployingNVIDIA DGX, HGX, GB-series AI systems, Spectrum-X, or Quantum InfiniBand.
- Knowledge ofNCCL, GPUDirect RDMA, and large-scale AI/ML workload traffic patterns, including MoE workloads.
- Familiarity withKubernetes networking, Slurm, parallel file systems, NVMe-oF, and related AI/HPC technologies.
- Experience withPython/Go, Ansible, CI/CD, and network automation.
- NVIDIA, Arista, or other relevant vendor certifications are an advantage.
3. Performance Indicators & Additional Requirements
3.1 Key Performance Indicators
- Deliver network projects on schedule with complete and accurate cabling, configuration, testing, and acceptance documentation.
- EnsureRDMA and NCCL performancemeets defined customer and production requirements.
- Respond rapidly to network incidents and accurately identify root causes, with complete RCA documentation and corrective action plans.
- Continuously improve network automation, operational efficiency, tenant security, and network isolation.
3.2 Additional Requirements
- Willingness to travel frequently acrossAsia-Pacificcountries.
- Able to support scheduled night-time maintenance, emergency troubleshooting, and on-call operations when required.
- Comfortable working in a fast-paced project environment with multiple stakeholders.
- Strong communication and coordination skills when working with data center operators, hardware vendors, enterprise customers, and internal engineering teams.
4. Job Summary
TheAI Cluster Network Engineeris a key technical role responsible for designing, deploying, optimizing, and operating high-performance networks forAI/GPU computing and GPUaaS environments.
The role is deeply focused on NVIDIA’s advanced networking technologies, includingSpectrum-X Ethernet, Quantum InfiniBand, RoCEv2, RDMA, and high-performance cluster networking. The successful candidate will combine strong traditional data center networking expertise with hands-on experience in AI/HPC networking and distributed workload optimization.
In addition to network architecture and performance tuning, the role coversmulti-tenant network isolation, automation, observability, production operations, customer POCs, and project deliveryacross the Asia-Pacific region.
This position is ideal for a senior network engineer who wants to specialize inAI infrastructure, GPU clusters, high-performance computing, and next-generation data center networking, with a strong focus on commercial, production-grade GPUaaS deployments.