Network Engineer (AI DC)

Vouch Recruitment

Singapore

On-site

SGD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Vouch Recruitment is assisting a fast-growing AI infrastructure company in Singapore. You will design, deploy, and support high-performance data center network infrastructure tailored for AI and GPU compute environments.

You will build Spine-Leaf architectures, deploy high-speed fabrics, and manage Layer 2/3 networking with BGP, OSPF, VXLAN EVPN, ECMP, and MLAG/VPC. Strong problem solving and collaboration across teams are essential.

Qualifications

  • Bachelor's degree in a relevant computing field is required.
  • 5+ years designing or supporting data center network infrastructure.
  • Strong grasp of modern data center networking principles.
  • Hands-on with routing/switching protocols: BGP, OSPF, VXLAN EVPN, ECMP, MLAG/VPC.
  • Experience with high-speed Ethernet networks (25G/40G/100G+).
  • Exposure to AI, HPC, GPU clusters or large-scale compute environments is a plus.
  • Familiarity with Cisco Nexus, Arista, Juniper, NVIDIA Spectrum or Mellanox hardware.
  • Good networking security concepts, Linux networking basics, and network troubleshooting.
  • Experience with Python, Ansible, REST APIs or similar automation is a plus.

Responsibilities

  • Design, deploy, and support high-performance data center network infrastructure for AI and GPU compute environments.
  • Build, configure, and optimize Spine-Leaf network architectures for large-scale GPU clusters.
  • Deploy and manage high-speed Ethernet and/or InfiniBand fabrics supporting AI/HPC workloads.
  • Configure Layer 2/3 networking (BGP, OSPF, VXLAN EVPN, ECMP, MLAG/VPC) and maintain routing.
  • Monitor, troubleshoot, and resolve complex network issues across compute, storage, and AI infrastructure.
  • Optimize network performance, latency, and throughput for distributed AI training and inference.
  • Perform firmware upgrades, capacity planning, and lifecycle management.
  • Collaborate with Platform, Infrastructure, Linux, DevOps, and AI Engineering teams.
  • Develop and maintain network documentation, procedures, and runbooks.
  • Implement monitoring, alerting, and automation to improve operations.
  • Participate in incident response and on-call production support.
  • Evaluate and recommend new networking technologies to improve scalability and reliability.

Skills

Network design
Troubleshooting
Cross-functional collaboration
Communication
On-call support

Education

Bachelor's Degree in Computer Science/Engineering/IT

Tools

Cisco Nexus
Arista
Juniper
NVIDIA Spectrum
Mellanox
InfiniBand
RoCEv2
RDMA
GPUDirect
SONiC
Cumulus Linux

Job description

Our client is a fast-growing AI infrastructure company building next-generation GPU-powered AI platforms. They operate high-performance AI data centers that support large-scale machine learning, AI model training and inference workloads. The environment is highly technical, focusing on low-latency networking, scalability, automation and operational excellence.

Primary Responsibilities

  • Design, deploy, and support high-performance data center network infrastructure for AI and GPU compute environments.
  • Build, configure, and optimize Spine-Leaf network architectures for large-scale GPU clusters.
  • Deploy and manage high-speed Ethernet and/or InfiniBand fabrics supporting AI/HPC workloads.
  • Configure and maintain Layer 2 and Layer 3 networking technologies, including BGP, OSPF, VXLAN EVPN, ECMP, and MLAG/VPC.
  • Monitor, troubleshoot, and resolve complex network issues across compute, storage, and AI infrastructure.
  • Optimize network performance, latency, and throughput to support distributed AI training and inference.
  • Perform firmware upgrades, network maintenance, capacity planning, and lifecycle management.
  • Collaborate closely with Platform, Infrastructure, Linux, DevOps, and AI Engineering teams to deliver scalable AI infrastructure.
  • Develop and maintain network documentation, operational procedures, and technical runbooks.
  • Implement network monitoring, alerting, and automation to improve operational efficiency.
  • Participate in incident response and on-call support for production environments.
  • Evaluate and recommend new networking technologies to improve scalability, reliability, and performance.

What We're Looking For

  • Bachelor's Degree in Computer Science, Computer Engineering, Information Technology, or a related discipline.
  • 5+ years of experience designing or supporting enterprise or data center network infrastructure.
  • Strong understanding of modern data center networking principles
  • Hands-on experience with routing and switching protocols such as BGP, OSPF, VXLAN EVPN, ECMP, and MLAG/VPC.
  • Experience managing high-speed Ethernet networks (25G/40G/100G/200G/400G).
  • Exposure to AI, HPC, GPU clusters, or large-scale compute environments would be highly advantageous.
  • Experience working with networking platforms such as Cisco Nexus, Arista, Juniper, NVIDIA Spectrum, or Mellanox.
  • Good understanding of network security concepts including segmentation, ACLs, and firewall policies.
  • Familiarity with Linux networking fundamentals and troubleshooting.
  • Experience with network automation using Python, Ansible, REST APIs, or similar tools is an advantage.
  • Knowledge of technologies such as InfiniBand, RoCEv2, RDMA, GPUDirect, SONiC, or Cumulus Linux is a plus.
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Excellent communication skills with the ability to collaborate effectively across cross-functional engineering teams.
  • Comfortable working in a fast-paced, high-availability production environment supporting mission-critical AI infrastructure.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network Engineer (Data Centre / GPU Infrastructure)
Network Engineer (Data Centre / GPU Infrastructure)

Visa Hunt • Singapore

On-site
SGD 90,000 - 140,000
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS
Senior Solutions Architect, Infiniband and Networking Ethernet - NVIS

NVIDIA Gruppe • Singapore

On-site
SGD 90,000 - 120,000
Hardware Engineer
Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Network Operations Engineer
Network Operations Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Senior Network Engineer
Senior Network Engineer

Equinix • Singapore

On-site
SGD 70,000 - 100,000
Senior Network Engineer, DGX Cloud - Backbone
Senior Network Engineer, DGX Cloud - Backbone

nvidia ai • Singapore

On-site
SGD 180,000 - 240,000
AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SwapeTech • Singapore

On-site
SGD 180,000 - 260,000
Hardware Engineer
Hardware Engineer

runsun cloud pte ltd • Singapore

On-site
SGD 100,000 - 160,000
Senior Network Engineer, AI Data Center (GPU)
Senior Network Engineer, AI Data Center (GPU)

Vouch Recruitment • Singapore

On-site
SGD 120,000 - 160,000
Network Engineer
Network Engineer

Firmus Technologies • Singapore

On-site
SGD 120,000 - 180,000