AI Infra Engineer: GPU Clusters & HPC Networking

The Supreme HR Advisory Pte Ltd

Singapore

On-site

SGD 56,000 - 78,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

The Supreme HR Advisory Pte Ltd seeks an AI Infrastructure Engineer to design, deploy and manage large multi‑GPU compute clusters in Singapore. You will optimize Linux systems, tune kernels, and automate with IaC to support distributed AI workloads.

You will work with Kubernetes, Slurm, Ray, and HPC schedulers; ensure low latency networking (RDMA/InfiniBand) and high‑IOPS storage; collaborate with AI/ML teams and maintain golden images and firmware updates.

Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, Computer Engineering, or equivalent practical experience.
  • 3–6+ years of hands-on experience in infrastructure engineering, HPC, DevOps, or cloud infrastructure.
  • Relevant certifications (e.g., CKA/CKAD, NVIDIA, AWS/Azure/GCP) are a plus.
  • Experience with multi-GPU compute clusters and AI/ML distributed workloads.

Responsibilities

  • Architect, configure, and maintain high-density multi-GPU compute clusters (NVIDIA HGX/DGX).
  • Implement and manage container orchestration platforms (Kubernetes, Slurm, or Ray) for AI workloads.
  • Monitor GPU health, telemetry, utilization, and thermals; minimize idle time and bottlenecks.

Skills

Linux administration
Kernel tuning
Shell scripting
Bash
Python
GPU architectures
CUDA runtimes
Kubernetes
Slurm
Ray
RDMA
InfiniBand
PFC configurations
High IO storage

Education

Bachelor’s degree in CS/IT/Engineering

Tools

Terraform
Ansible
Helm
Pulumi

Job description

The Supreme HR Advisory Pte Ltd seeks an AI Infrastructure Engineer to design, deploy and manage large multi‑GPU compute clusters in Singapore. You will optimize Linux systems, tune kernels, and automate with IaC to support distributed AI workloads.

You will work with Kubernetes, Slurm, Ray, and HPC schedulers; ensure low latency networking (RDMA/InfiniBand) and high‑IOPS storage; collaborate with AI/ML teams and maintain golden images and firmware updates.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Compute & HPC Infrastructure Engineer
Senior AI Compute & HPC Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
Senior AI Infra Engineer — GPUs, HPC & Kubernetes
Senior AI Infra Engineer — GPUs, HPC & Kubernetes

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra Architect: GPU Clusters & HPC
AI Infra Architect: GPU Clusters & HPC

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000
AI Infra Engineer: HPC Clusters & AI Workloads
AI Infra Engineer: HPC Clusters & AI Workloads

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra Architect: GPU Clusters & HPC
AI Infra Architect: GPU Clusters & HPC

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra DevOps Engineer - GPU/Cloud HPC
AI Infra DevOps Engineer - GPU/Cloud HPC

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 57,000 - 77,000
AI Infrastructure Engineer - GPU HPC & Kubernetes Expert
AI Infrastructure Engineer - GPU HPC & Kubernetes Expert

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Systems Infra Engineer - Multi-GPU HPC
AI Systems Infra Engineer - Multi-GPU HPC

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
GPU AI Infra Engineer for High-Perf Clusters
GPU AI Infra Engineer for High-Perf Clusters

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
GPU HPC Infrastructure Engineer
GPU HPC Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
5-day work week