Senior Network Engineer

Trulyyy

Singapore

On-site

SGD 120,000 - 180,000

Full time

11 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Trulyyy is seeking an experienced AI Cloud Network Operations Engineer in Singapore to operate and optimize high-performance AI/HPC network infrastructure across data centers. The role emphasizes incident handling, performance troubleshooting, and scalable network configuration for large GPU clusters.

Responsibilities include AI network operations, fabric optimization with InfiniBand/RoCEv2, and automation using Python/Go; familiarity with Prometheus/Grafana/Zabbix is expected.

Qualifications

  • 5+ years of Network Operations / Network Engineering experience in large-scale Cloud, Internet, Data Center, Carrier or HPC environments.
  • Hands-on knowledge of BGP, OSPF, VXLAN, EVPN and ECMP with independent troubleshooting on enterprise networking equipment.
  • Exposure to InfiniBand and/or RoCEv2 with understanding of PFC and ECN for high-performance workloads.
  • Experience with monitoring tools such as Prometheus, Grafana, Zabbix and related platforms.

Responsibilities

  • AI Network Operations — Monitor and operate large-scale data center and AI/HPC networks across switches, routers, links, bandwidth and health.
  • Incident & Performance Troubleshooting — Own network incidents and troubleshoot latency, jitter, packet loss and NCCL throughput issues.
  • Network Configuration & Changes — Execute and troubleshoot BGP, OSPF, VXLAN, EVPN and ECMP environments, VLANs and bandwidth changes.
  • AI/HPC Fabric Operations — Support GPU cluster networking using InfiniBand and/or RoCEv2 with lossless Ethernet tech like PFC and ECN.
  • Monitoring & Automation — Improve visibility and efficiency using monitoring platforms and Python/Go-based automation.

Skills

Network Operations
BGP
OSPF
VXLAN
EVPN
ECMP
InfiniBand
RoCEv2
NVIDIA Spectrum/Quantum
Prometheus
Grafana
Zabbix
Python
Go

Tools

Prometheus
Grafana
Zabbix
NVIDIA Spectrum/Quantum
Arista
Cisco

Job description

Our client is a global technology company operating large-scale AI/HPC and data center infrastructure across multiple international markets. As its AI cloud capabilities continue to scale, the company is looking for an experienced AI Cloud Network Operations Engineer to operate and optimize high-performance network infrastructure supporting large-scale GPU computing environments.

Job Responsibilities
  • AI Network Operations — Monitor and operate large-scale data center and AI/HPC networks across switches, routers, optical links, bandwidth utilization and network health.
  • Incident & Performance Troubleshooting — Own network incidents and troubleshoot complex performance issues including latency/jitter, packet loss, GPU-to-GPU communication and NCCL throughput degradation.
  • Network Configuration & Changes — Execute and troubleshoot BGP, OSPF, VXLAN, EVPN and ECMP environments, including VLAN, routing, traffic isolation and bandwidth changes.
  • AI/HPC Fabric Operations — Support high-performance GPU cluster networking using InfiniBand and/or RoCEv2, including lossless Ethernet technologies such as PFC and ECN.
  • Monitoring & Automation — Improve network visibility, operational efficiency and reliability using monitoring platforms and Python/Go-based network automation.
Job Requirements
  • 5+ years of Network Operations / Network Engineering experience within large-scale Cloud, Internet, Data Center, Carrier or HPC environments.
  • Strong hands-on knowledge of BGP, OSPF, VXLAN, EVPN and ECMP, with independent troubleshooting capability on enterprise networking equipment.
  • Practical exposure to InfiniBand and/or RoCEv2, with understanding of PFC, ECN and lossless networking for high-performance workloads.
  • Experience operating large-scale infrastructure using vendors such as NVIDIA Spectrum/Quantum, Arista or Cisco, together with monitoring tools such as Prometheus, Grafana, Zabbix or equivalent.
  • Exposure to large-scale GPU clusters / AI infrastructure is highly advantageous, particularly NVIDIA H100/GB200 environments, NCCL/MPI, NVIDIA UFM, network automation or optical/DCI networks.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Network Engineer – GPU / AI Infrastructure
Network Engineer – GPU / AI Infrastructure

VOUCH RECRUITMENT PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Network Operations Engineer
Network Operations Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Network Engineer, GPUaaS
Network Engineer, GPUaaS

Singtel • Singapore

On-site
Confidential
Network Engineer
Network Engineer

SUPERPOWER X AI (SINGAPORE) TECHNOLOGY PTE. LTD. • Singapore

On-site
SGD 70,000 - 120,000
Hardware Engineer
Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Network Engineer – AI Network & Security
Network Engineer – AI Network & Security

Firmus Technologies • Singapore

On-site
SGD 120,000 - 170,000
Diversity commitment
AI Infrastructure Network Engineer
AI Infrastructure Network Engineer

Tencent • Singapore

On-site
SGD 180,000 - 240,000
AI Cloud Network Engineer for HPC & GPU Clusters
AI Cloud Network Engineer for HPC & GPU Clusters

Trulyyy • Singapore

On-site
SGD 120,000 - 180,000
System Engineer
System Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Infrastructure Support Engineer
Senior AI Infrastructure Support Engineer

nscale operations apac pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000