AI Cloud Network Engineer for HPC & GPU Clusters

Trulyyy

Singapore

On-site

SGD 120,000 - 180,000

Full time

11 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Trulyyy is seeking an experienced AI Cloud Network Operations Engineer in Singapore to operate and optimize high-performance AI/HPC network infrastructure across data centers. The role emphasizes incident handling, performance troubleshooting, and scalable network configuration for large GPU clusters.

Responsibilities include AI network operations, fabric optimization with InfiniBand/RoCEv2, and automation using Python/Go; familiarity with Prometheus/Grafana/Zabbix is expected.

Qualifications

  • 5+ years of Network Operations / Network Engineering experience in large-scale Cloud, Internet, Data Center, Carrier or HPC environments.
  • Hands-on knowledge of BGP, OSPF, VXLAN, EVPN and ECMP with independent troubleshooting on enterprise networking equipment.
  • Exposure to InfiniBand and/or RoCEv2 with understanding of PFC and ECN for high-performance workloads.
  • Experience with monitoring tools such as Prometheus, Grafana, Zabbix and related platforms.

Responsibilities

  • AI Network Operations — Monitor and operate large-scale data center and AI/HPC networks across switches, routers, links, bandwidth and health.
  • Incident & Performance Troubleshooting — Own network incidents and troubleshoot latency, jitter, packet loss and NCCL throughput issues.
  • Network Configuration & Changes — Execute and troubleshoot BGP, OSPF, VXLAN, EVPN and ECMP environments, VLANs and bandwidth changes.
  • AI/HPC Fabric Operations — Support GPU cluster networking using InfiniBand and/or RoCEv2 with lossless Ethernet tech like PFC and ECN.
  • Monitoring & Automation — Improve visibility and efficiency using monitoring platforms and Python/Go-based automation.

Skills

Network Operations
BGP
OSPF
VXLAN
EVPN
ECMP
InfiniBand
RoCEv2
NVIDIA Spectrum/Quantum
Prometheus
Grafana
Zabbix
Python
Go

Tools

Prometheus
Grafana
Zabbix
NVIDIA Spectrum/Quantum
Arista
Cisco

Job description

Trulyyy is seeking an experienced AI Cloud Network Operations Engineer in Singapore to operate and optimize high-performance AI/HPC network infrastructure across data centers. The role emphasizes incident handling, performance troubleshooting, and scalable network configuration for large GPU clusters.

Responsibilities include AI network operations, fabric optimization with InfiniBand/RoCEv2, and automation using Python/Go; familiarity with Prometheus/Grafana/Zabbix is expected.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infra Engineer: GPU Clusters & HPC Networking
AI Infra Engineer: GPU Clusters & HPC Networking

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000
Senior AI Compute & HPC Infrastructure Engineer
Senior AI Compute & HPC Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
Senior AI Infra Architect: GPU Clusters & HPC Ops
Senior AI Infra Architect: GPU Clusters & HPC Ops

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra DevOps Engineer - GPU/Cloud HPC
AI Infra DevOps Engineer - GPU/Cloud HPC

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 57,000 - 77,000
AI Infra Engineer: HPC GPU Clusters & Kubernetes
AI Infra Engineer: HPC GPU Clusters & Kubernetes

The Supreme HR Advisory Pte. Ltd. • Singapore

On-site
SGD 56,000 - 78,000
GPU HPC Infra Engineer for AI Clusters
GPU HPC Infra Engineer for AI Clusters

The Supreme HR Advisory Pte. Ltd. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer — GPU & HPC Clusters
AI Infrastructure Engineer — GPU & HPC Clusters

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
HPC/AI Cluster Engineer (GPU Specialist)
HPC/AI Cluster Engineer (GPU Specialist)

PaleBlueDot AI • Singapore

On-site
SGD 120,000 - 180,000
GPU-Accelerated AI Infra Network Engineer
GPU-Accelerated AI Infra Network Engineer

Singtel • Singapore

On-site
Confidential
AI HPC Infra Engineer — GPU Clusters & Slurm Expert
AI HPC Infra Engineer — GPU Clusters & Slurm Expert

The Supreme HR Advisory Pte. Ltd. • Singapore

On-site
SGD 56,000 - 78,000