Senior AI Infrastructure & Networking Engineer

Genesis Networks Pte Ltd

Singapore

On-site

SGD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Genesis Networks Pte Ltd is seeking an expert Senior AI Infrastructure & Networking Engineer to lead the architecture, deployment, and optimization of our AI Factory. You will design, deploy, and scale GPU clusters up to 512+ nodes with NVIDIA Blackwell Ultra systems.

You will bridge physical infrastructure with advanced fabrics, ensure lossless transport, and drive automation using Ansible and Terraform in Kubernetes environments.

Qualifications

  • Bachelor’s or Master’s degree in CS, network engineering, or related field.
  • Proven AI networking experience with RoCE v2 and traffic optimization.
  • Experience with high-density scale-up/scale-out systems and NVIDIA HGX/DGX.
  • Experience with cluster deployment suites and MLOps frameworks.
  • Strong datacenter routing knowledge and EVPN/VXLAN experience.
  • Layer-1 validation, fiber cabling, and diagnostics expertise.

Responsibilities

  • Design, build, and optimize high-throughput AI compute networks.
  • Configure RoCE v2, PFC, ECN, and DCQCN for low latency.
  • Implement fat-tree topologies for 8-GPU servers to maximize NCCL perf.
  • Deploy and manage high-speed interfaces (ConnectX-8, BlueField-3 DPUs).
  • Drive IaC deployments with Ansible and Terraform within Kubernetes.
  • Use Telemetry tools (NetQ, WJH) for real-time diagnostics and benchmarking.
  • Collaborate with data center facilities on high-density environment metrics.

Skills

AI networking
RoCE v2
High-density HPC
NCCL
Cluster management
Routing protocols
eBGP IPv6

Education

Bachelor’s or Master’s degree

Tools

NVIDIA Mission Control
Base Command Manager
Run:ai
NVIDIA Network Operator

Job description

We are seeking an expert Senior AI Infrastructure & Networking Engineer to lead the architecture, deployment, and optimization of our next-generation AI Factory. In this role, you will be responsible for building and scaling high-density GPU supercomputing clusters (up to 512+ nodes) featuring NVIDIA Blackwell Ultra B300 systems. You will bridge the gap between heavy physical infrastructure (liquid cooling/busbar power) and advanced logical fabrics, ensuring predictable, line-rate, and lossless transport for massive generative AI training and reasoning workloads.

Key Responsibilities
  • AI Fabric Architecture & Deployment: Design, build, and optimize high-throughput, ultra-low-latency East-West compute networks using NVIDIA Spectrum-X Ethernet platforms (Spectrum-4 ASICs) and/or NVIDIA Quantum-X800 InfiniBand switching
  • Performance Tuning for Lossless Networking: Configure and fine-tune critical Layer 2/3 lossless transport mechanisms, including Remote Direct Memory Access over Converged Ethernet (RoCE v2), Priority Flow Control (PFC), Explicit Congestion Notification (ECN), and DCQCN
  • Rail-Optimized Topologies: Implement and maintain non-blocking, multi-plane, full fat-tree network topologies mapped to 8-GPU server architectures to maximize collective communication performance via NCCL (NVIDIA Collective Communications Library)
  • SmartNIC & DPU Management: Deploy and manage high-speed compute network interfaces, including ConnectX-8 SuperNICs (800 Gb/s) and BlueField-3 DPUs for isolated infrastructure management, storage acceleration, and multi-tenant security
  • Full-Stack Orchestration & Automation: Drive infrastructure-as-code deployments using Ansible and Terraform. Initialize and monitor the NVIDIA Network Operator within core Kubernetes orchestration layers
  • Telemetry & Validation: Utilize deep network telemetry tools such as NVIDIA NetQ and "What Just Happened" (WJH) to stream real-time switch diagnostics. Conduct line-rate cluster benchmarking using ib_write_bw and ib_write_lat to eliminate physical layer bottlenecks
  • Cross-Functional Infrastructure Alignment: Collaborate closely with data center facility teams on high-density environment metrics (~15–20 kW+ per rack, liquid-cooled rows, Coolant Distribution Units (CDUs), and Rear Door Heat Exchangers). Ensure operational verification aligns with international standards (e.g., IDCA G-Grade or Uptime Institute)
Required Technical Skills & Qualifications
  • Education: Bachelor’s or Master’s degree in Computer Science, Network Engineering, Systems Engineering, or a related technical discipline
  • AI Networking Expertise: Proven track record of configuring RoCE v2, adaptive routing, and traffic optimization specifically for machine learning/HPC workloads
  • Hardware Familiarity: Deep understanding of high-density scale-up and scale-out systems (NVIDIA HGX/DGX architectures, PCIe switching, OSFP/QSFP112 optical and copper assemblies)
  • Software & Cluster Management: Experience with cluster deployment suites like NVIDIA Mission Control, Base Command Manager, Run:ai, or similar enterprise MLOps frameworks
  • Routing Protocols: Strong proficiency with advanced datacenter networking protocols, particularly eBGP IPv6 unnumbered underlays and EVPN/VXLAN overlays for multi-tenant isolation
  • Cabling & Layer 1 Validation: Experience managing complex structured fiber trunking (MPO-12/MPO-24 APC) and executing layer-1 diagnostics (ibdiagnet, iblinkinfo)
Preferred Certifications
  • NVIDIA Certified Professional - AI Networking (NCP-AIN) (Highly Preferred)
  • NVIDIA Certified Expert - Cloud End-to-End Fabric (NCE-CEF)
  • Advanced networking tracks from major vendors (e.g., CCIE, JNCIE, or Nokia Service Routing Architect) combined with proven data center fabric experience
What We Offer
  • Opportunity to work with first-of-its-kind, world-class AI supercomputing technologies (NVIDIA Blackwell Ultra)
  • High-impact role shaping the foundational architecture for enterprise generative AI and large-scale LLM initiatives
  • Competitive salary, comprehensive benefits package, and continuous learning paths for advanced AI operations certifications
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network Engineer (AI DC)
Network Engineer (AI DC)

Vouch Recruitment • Singapore

On-site
SGD 120,000 - 160,000
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS

NVIDIA Gruppe • Singapore

On-site
SGD 90,000 - 120,000
Hardware Engineer
Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Infrastructure & Networking Architect
Senior AI Infrastructure & Networking Architect

Genesis Networks Pte Ltd • Singapore

On-site
SGD 180,000 - 240,000
AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SwapeTech • Singapore

On-site
SGD 180,000 - 260,000
Network Engineer (Data Centre / GPU Infrastructure)
Network Engineer (Data Centre / GPU Infrastructure)

Visa Hunt • Singapore

On-site
SGD 90,000 - 140,000
HPC Engineer
HPC Engineer

NVIDIA Gruppe • Singapore

On-site
SGD 120,000 - 180,000
HPC Engineer
HPC Engineer

NVIDIA • Singapore

On-site
SGD 120,000 - 180,000
AI Infrastructure Engineer
AI Infrastructure Engineer

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000
Senior Network Engineer (GPU-as-a-Service Infrastructure)
Senior Network Engineer (GPU-as-a-Service Infrastructure)

BGC Group Pte Ltd • Singapore

On-site
SGD 50,000 - 80,000