AI Infrastructure Engineer — GPU & HPC Clusters

RUNSUN SERVICE PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

RUNSUN SERVICE PTE. LTD. in Singapore seeks an experienced AI System Engineer to design, deploy, operate, and optimize AI training clusters and GPU platforms. The role involves managing Linux systems, CUDA/NVIDIA drivers, and container orchestration to ensure high availability.

You will build and maintain scalable AI infrastructure, automate operations, and work with datacenter teams to resolve training environment issues, while supporting on-call rotations and occasional travel as required.

Qualifications

  • Bachelor's degree or above in Computer Engineering, Electrical Engineering, Telecommunications, or related fields.
  • 3+ years of Linux administration experience; strong knowledge of Ubuntu, Rocky Linux, and RHEL; familiarity with boot process, kernel, filesystems, and performance tuning.
  • Ability to troubleshoot complex system issues independently.
  • Experience with NVIDIA GPU products H100, H200, B200, B300, GB200 NVL72 and GB300 NVL72.
  • Familiar with CUDA, NCCL, NV Link, NV Switch, GPU Direct RDMA.
  • Understanding of distributed AI training architectures.
  • Hands-on experience with Kubernetes; familiarity with Docker and Containerd; Helm knowledge.

Responsibilities

  • Deploy and operate AI training and HPC clusters.
  • Install, configure, and optimize operating systems on GPU servers.
  • Manage cluster resources and capacity.
  • Develop infrastructure automation tools and scripts.
  • Build monitoring and observability platforms and ensure high availability.

Skills

Linux administration
GPU computing
Container orchestration
Scripting (Shell, Python)
Network troubleshooting

Education

Bachelor's degree in Computer/Electrical Engineering or related fields

Tools

Kubernetes
Docker
Containerd
NVIDIA CUDA
NCCL
NVIDIA drivers

Job description

RUNSUN SERVICE PTE. LTD. in Singapore seeks an experienced AI System Engineer to design, deploy, operate, and optimize AI training clusters and GPU platforms. The role involves managing Linux systems, CUDA/NVIDIA drivers, and container orchestration to ensure high availability.

You will build and maintain scalable AI infrastructure, automate operations, and work with datacenter teams to resolve training environment issues, while supporting on-call rotations and occasional travel as required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Training Cluster Hardware Engineer
AI Training Cluster Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Compute & HPC Infrastructure Engineer
Senior AI Compute & HPC Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
Senior AI Infra Architect: GPU Clusters & HPC Ops
Senior AI Infra Architect: GPU Clusters & HPC Ops

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
Staff AI Infrastructure Systems Engineer
Staff AI Infrastructure Systems Engineer

Cloudera • Singapore

On-site
SGD 180,000 - 260,000
Generous PTO Policy
Flexible WFH Policy
Mental & Physical Wellness programs
+2
AI Infra Engineer: HPC GPU Clusters & Kubernetes
AI Infra Engineer: HPC GPU Clusters & Kubernetes

The Supreme HR Advisory Pte. Ltd. • Singapore

On-site
SGD 56,000 - 78,000
GPU HPC Infra Engineer for AI Clusters
GPU HPC Infra Engineer for AI Clusters

The Supreme HR Advisory Pte. Ltd. • Singapore

On-site
SGD 56,000 - 78,000
Senior AI Compute Architect - HPC Infra Lead
Senior AI Compute Architect - HPC Infra Lead

NVIDIA Corporation • Singapore

On-site
SGD 180,000 - 260,000
GPU AI Infrastructure Engineer II
GPU AI Infrastructure Engineer II

Proxima Beta Pte. Limited • Singapore

On-site
SGD 120,000 - 180,000
AI Systems Infra Engineer - Multi-GPU HPC
AI Systems Infra Engineer - Multi-GPU HPC

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
HPC/AI Cluster Engineer (GPU Specialist)
HPC/AI Cluster Engineer (GPU Specialist)

PaleBlueDot AI • Singapore

On-site
SGD 120,000 - 180,000