AI Infrastructure Network Engineer

Tencent

Singapore

On-site

SGD 180,000 - 240,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Tencent seeks an experienced AI Infrastructure Network Engineer / Architect to plan, deliver, and operate large-scale GPU cluster networks for internal AI teams. You will bridge internal AI demand with external vendors across regions, validating designs and ensuring production readiness.

The role requires strong data center networking, GPU interconnect expertise, and cross‑team coordination to deliver scalable, reliable AI compute clusters.

Qualifications

  • Bachelor's degree or higher in a relevant technical field.
  • 5+ years in data center networking, cloud infra, or large-scale delivery.
  • Solid AI infrastructure understanding for distributed training workloads.
  • Familiar with InfiniBand and RoCEv2 in data center contexts.
  • Experience with vendor reviews, design validation, and production readiness.

Responsibilities

  • Translate business/workload requirements into GPU cluster network specs.
  • Review and validate GPU cluster network architecture and plans.
  • Engage with external vendors and system integrators for delivery.
  • Monitor delivery progress, identify risks, coordinate corrective actions.
  • Define acceptance criteria and support performance validation.
  • Coordinate handover to internal teams with docs and procedures.

Skills

Data center networking
GPU cluster networking
InfiniBand networking
RoCEv2
Spine-Leaf topology
Vendor management
Network troubleshooting
Capacity planning
Cross-team collab
Linux networking

Education

Bachelor's degree in CS/EE/IT

Tools

InfiniBand
RoCEv2
VXLAN
Spine-Leaf
Monitoring tools
TCP/IP
DNS
Firewall rules

Job description

We are looking for an experienced AI Infrastructure Network Engineer / Architect to support the planning, delivery, acceptance, and ongoing operations of large-scale GPU cluster infrastructure for our internal AI business teams.

The AI Compute Centre sits within Tencent's Overseas IT department, acting as the bridge between internal AI infrastructure demand and the external resources that fulfill it. We own the full lifecycle of AI compute clusters — requirement gathering, capacity planning, architecture review, delivery coordination, and day-to-day operations — across regions worldwide.

In this role, you will act as the key technical bridge between internal AI business stakeholders and external GPU cluster vendors, network equipment vendors, colocation partners, and managed infrastructure providers. You will be responsible for understanding business and technical requirements, translating them into clear infrastructure and network requirements, overseeing vendor design and delivery, validating the final implementation, and ensuring the cluster is successfully handed over to business users and operated reliably in production.

This role is ideal for someone who has strong data center networking and AI infrastructure knowledge, hands-on experience with GPU cluster networks, and the ability to coordinate across business teams, engineering teams, and external vendors.

Key Responsibilities

  • Work closely with internal AI business teams, AI platform teams, and infrastructure stakeholders to understand AI workload requirements, including training, inference, data access, storage, bandwidth, latency, scalability, and reliability requirements.
  • Translate business and platform requirements into clear technical requirements for GPU cluster infrastructure, especially networking architecture, interconnect design, capacity planning, and operational requirements.
  • Engage with external vendors, including GPU cluster providers, network equipment vendors, colocation providers, and system integrators, to review proposed solutions and ensure they meet business and technical expectations.
  • Review and validate GPU cluster network architecture, including InfiniBand / RoCE networks, Spine-Leaf / Clos topologies, front-end service networks, storage networks, and out-of-band management networks.
  • Participate in the planning of foundational network resources, including IP addressing, VLAN/VXLAN, routing, bandwidth capacity, network segmentation, and management access.
  • Review vendor-provided architecture documents, topology diagrams, implementation plans, configuration standards, test plans, and operational runbooks.
  • Monitor and drive vendor delivery progress, identify technical risks or delivery gaps, and coordinate corrective actions to ensure the cluster is delivered on time and according to requirements.
  • Define and execute acceptance criteria for GPU cluster delivery, including network connectivity, bandwidth, latency, redundancy, congestion control, fault tolerance, GPU-to-GPU communication performance, and cluster-level benchmark validation.
  • Support performance validation and troubleshooting for AI training and inference workloads, including network-related issues affecting distributed training, collective communication, storage access, or application performance.
  • Coordinate the handover of validated GPU clusters to internal business and platform teams, including documentation, knowledge transfer, operational procedures, and post-handover support.
  • Own or support daily operations and maintenance of AI infrastructure networks, including incident response, troubleshooting, configuration review, capacity monitoring, change management, and continuous improvement.
  • Develop and maintain technical documentation, including HLD/LLD, network topology diagrams, configuration baselines, acceptance reports, SOPs, and operational playbooks.

Qualifications

  • Bachelor's degree or above in Computer Science, Telecommunications, Electrical Engineering, Information Technology, or a related technical field.
  • 5+ years of experience in data center networking, cloud infrastructure networking, backbone networking, or large-scale infrastructure delivery.
  • Solid understanding of AI infrastructure and GPU cluster networking requirements, especially for distributed training and high-performance computing workloads.
  • Familiarity with high-performance GPU cluster interconnect technologies such as InfiniBand and/or RoCEv2.
  • Good understanding of AI cluster components, including GPU servers, high-speed networking, storage systems, management networks, and cluster orchestration platforms.
  • Experience working with infrastructure vendors, system integrators, colocation providers, or cloud infrastructure providers.
  • Ability to review and challenge vendor designs, implementation plans, test results, and operational documents from both architecture and production-readiness perspectives.
  • Experience defining or executing infrastructure acceptance tests, including network performance, redundancy, failure recovery, connectivity, and stability validation.
  • Strong troubleshooting skills across L2/L3 networking, Linux networking, TCP/IP, routing, DNS, firewall/security rules, and data center connectivity.
  • Strong project coordination skills, with the ability to track vendor delivery progress, identify risks, drive issue resolution, and communicate clearly with both technical and non-technical stakeholders.
  • Excellent written and verbal communication skills, with the ability to translate business needs into technical requirements and explain technical trade-offs to business teams.
  • Self-motivated, responsible, detail-oriented, and comfortable working in a fast-paced environment with multiple internal and external stakeholders.
  • Customer-oriented mindset and strong ownership in ensuring business teams receive stable, performant, and production-ready AI infrastructure.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Network Engineer – GPU / AI Infrastructure
Network Engineer – GPU / AI Infrastructure

VOUCH RECRUITMENT PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI GPU Compute Network Architect
AI GPU Compute Network Architect

Tencent • Singapore

On-site
SGD 180,000 - 240,000
Network Operations Engineer
Network Operations Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Network Engineer, GPUaaS
Network Engineer, GPUaaS

Singtel • Singapore

On-site
Confidential
Network Engineer, GPUaaS (Singapore, Singapore)
Network Engineer, GPUaaS (Singapore, Singapore)

Singtel • Singapore

On-site
Confidential
Hardware Engineer
Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
System Engineer
System Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Global Head of AI Compute Infrastructure
Global Head of AI Compute Infrastructure

Kuailu Software • Singapore

On-site
SGD 220,000 - 380,000
Global Head of AI Compute Infrastructure
Global Head of AI Compute Infrastructure

UMELIFE (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 250,000 - 400,000
Global Head of AI Compute Infrastructure
Global Head of AI Compute Infrastructure

KUAILU SOFTWARE (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 250,000 - 380,000