AI Systems Engineer: GPU & HPC Clusters

runsun cloud pte ltd

Singapore

On-site

SGD 120,000 - 180,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

RUNSUN CLOUD PTE LTD is seeking an experienced AI System Engineer to design, deploy, operate, and optimize AI training clusters, GPU computing platforms, and supporting infrastructure in Singapore. The role emphasizes Linux expertise, GPU environments, containers, and HPC architectures.

Responsibilities include deploying AI training clusters, managing NVIDIA GPUs, CUDA drivers, and Fabric Manager, plus Kubernetes orchestration, automation, and platform development.

Qualifications

  • Bachelor's degree or above in Computer Engineering, Electrical Engineering, Telecommunications, or related fields.
  • 3+ years of Linux administration experience; strong knowledge of Ubuntu, Rocky Linux, and RHEL; troubleshooting skills.
  • Familiarity with CUDA, NCCL, NV Link, NV Switch, GPU Direct RDMA; understanding of distributed AI training architectures.
  • Experience with NVIDIA GPU products H100, H200, B200, B300 and related tools (Fabric Manager, drivers).
  • Hands-on Kubernetes, Docker, Helm; GPU scheduling and container runtimes; scripting in Shell and Python.

Responsibilities

  • Deploy and operate AI training and HPC clusters.
  • Install, configure, and optimize operating systems on GPU servers.
  • Manage cluster resources and capacity.
  • Perform system upgrades, patch management, and change implementation.
  • Develop and maintain standardized operational procedures.
  • Manage large-scale Linux environments and perform system performance tuning.
  • Analyze system logs and kernel issues; troubleshoot stability problems.
  • Manage user access and security policies.
  • Manage NVIDIA GPU platforms; deploy CUDA, drivers, and Fabric Manager; troubleshoot GPU issues.
  • Build and maintain Kubernetes clusters; support AI workload scheduling; ensure high availability.
  • Monitor infrastructure health and performance; automate operations; improve efficiency.

Skills

Linux systems
GPU computing
Kubernetes
NVIDIA GPUs
CUDA/NVLink
Shell/Python
Docker
Networking
Storage
Automation tooling

Education

Bachelor's degree or above in Computer/Electrical/Telecommunications or related field

Tools

Kubernetes
Docker
Helm
NVIDIA CUDA drivers
NVIDIA Base Command Manager

Job description

RUNSUN CLOUD PTE LTD is seeking an experienced AI System Engineer to design, deploy, operate, and optimize AI training clusters, GPU computing platforms, and supporting infrastructure in Singapore. The role emphasizes Linux expertise, GPU environments, containers, and HPC architectures.

Responsibilities include deploying AI training clusters, managing NVIDIA GPUs, CUDA drivers, and Fabric Manager, plus Kubernetes orchestration, automation, and platform development.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer — GPU & HPC Clusters
AI Infrastructure Engineer — GPU & HPC Clusters

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI GPU Cluster Engineer
AI GPU Cluster Engineer

runsun cloud pte ltd • Singapore

On-site
SGD 100,000 - 160,000
AI Training Cluster Hardware Engineer
AI Training Cluster Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior AI/HPC Systems Engineer
Senior AI/HPC Systems Engineer

NVIDIA • Singapore

On-site
SGD 120,000 - 180,000
HPC AI Lead Engineer — Cloud & GPU Cluster Architect
HPC AI Lead Engineer — Cloud & GPU Cluster Architect

National University Of Singapore • Singapore

On-site
SGD 80,000 - 120,000
GPU Data Centre Operations Engineer for AI & HPC
GPU Data Centre Operations Engineer for AI & HPC

Singtel Group • Singapore

On-site
SGD 50,000 - 80,000
Health and wellness benefits
Ongoing training and development programs
Internal mobility opportunities
AI Infrastructure Engineer - Large-Scale GPU Clusters
AI Infrastructure Engineer - Large-Scale GPU Clusters

SwapeTech • Singapore

On-site
SGD 180,000 - 260,000
GPU Cloud & AI Infrastructure Lead
GPU Cloud & AI Infrastructure Lead

zy future international pte. ltd. • Singapore

On-site
SGD 167,000 - 223,000
Lead GPU Data Center Engineer – AI Infrastructure
Lead GPU Data Center Engineer – AI Infrastructure

Hamilton Barnes Associates Limited • Singapore

On-site
SGD 160,000 - 220,000
10% bonus
Stock options
Hardware Engineer
Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000