GPU HPC Infra Engineer for AI Clusters

The Supreme HR Advisory Pte. Ltd.

Singapore

On-site

SGD 56,000 - 78,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

The Supreme HR Advisory Pte. Ltd. is seeking an AI Infrastructure Engineer in Singapore to architect and maintain high-density GPU compute clusters, optimize distributed AI workloads, and drive automation across deployment pipelines.

You will pair with AI/ML teams to monitor performance, scale storage, and ensure low-latency networking for distributed training. Minimum 3–6+ years in infra engineering, HPC, or DevOps is expected, with strong Linux, GPU, and Kubernetes expertise.

Qualifications

  • Linux systems administration, kernel tuning, and shell scripting (Bash/Python).
  • GPU hardware architectures, CUDA runtimes, and PCIe/NVLink knowledge.
  • Kubernetes GPU operator and HPC schedulers (Slurm, Run:AI, Ray).
  • High-speed networking: RDMA, RoCE v2, InfiniBand; storage for AI data.
  • IaC: Terraform and configuration management with Ansible.
  • Bachelor’s degree in CS/IT/Engineering or equivalent.
  • 3–6+ years of hands‑on infra engineering, HPC, DevOps, or cloud infra.
  • Certifications such as CKA/CKAD, NVIDIA certs, AWS/Azure/GCP are a plus.

Responsibilities

  • Compute & Cluster Management: Architect, configure, and maintain high-density multi-GPU compute clusters (NVIDIA HGX/DGX).
  • Container orchestration: Implement and manage Kubernetes, Slurm, or Ray optimized for AI/ML workloads.
  • Monitor GPU health, telemetry, utilization, and thermals; minimize idle compute time and prevent bottlenecks.
  • High-Performance Networking & Storage: Design and optimize low-latency fabrics and scalable storage (Lustre, Ceph, MinIO).
  • Automation & IaC: Build and manage deployment pipelines using Terraform, Ansible, Helm, or Pulumi; maintain golden images and kernel tuning.
  • Operations, Observability & Performance: Set up monitoring dashboards (Prometheus, Grafana, DCGM); diagnose bottlenecks; lead RCA and DR activities.

Skills

Linux systems administration
Kernel tuning
Shell scripting (Bash/Python)
GPU compute understanding
Distributed computing concepts
Observability
Incident response
NVIDIA GPU internals

Education

Bachelor’s Degree in Computer Science, Information Technology, Computer Engineering, or equivalent

Tools

Kubernetes
Slurm
Ray
Terraform
Ansible
Helm
Pulumi
Prometheus
Grafana
NVIDIA SMI/DCGM

Job description

The Supreme HR Advisory Pte. Ltd. is seeking an AI Infrastructure Engineer in Singapore to architect and maintain high-density GPU compute clusters, optimize distributed AI workloads, and drive automation across deployment pipelines.

You will pair with AI/ML teams to monitor performance, scale storage, and ensure low-latency networking for distributed training. Minimum 3–6+ years in infra engineering, HPC, or DevOps is expected, with strong Linux, GPU, and Kubernetes expertise.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Infra Architect: GPU Clusters & HPC Ops
Senior AI Infra Architect: GPU Clusters & HPC Ops

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
GPU HPC Infrastructure Engineer
GPU HPC Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
5-day work week
AI Infra Engineer: HPC GPU Clusters & Kubernetes
AI Infra Engineer: HPC GPU Clusters & Kubernetes

The Supreme HR Advisory Pte. Ltd. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra Engineer: GPU Clusters & HPC Networking
AI Infra Engineer: GPU Clusters & HPC Networking

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000
GPU AI Infra Engineer for High-Perf Clusters
GPU AI Infra Engineer for High-Perf Clusters

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
Senior AI Compute & HPC Infrastructure Engineer
Senior AI Compute & HPC Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI HPC Infra Engineer — GPU Clusters & Slurm Expert
AI HPC Infra Engineer — GPU Clusters & Slurm Expert

The Supreme HR Advisory Pte. Ltd. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra DevOps Engineer - GPU/Cloud HPC
AI Infra DevOps Engineer - GPU/Cloud HPC

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 57,000 - 77,000
AI Infra Engineer: GPU HPC Clusters & Orchestration
AI Infra Engineer: GPU HPC Clusters & Orchestration

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra Architect: GPU Clusters & HPC
AI Infra Architect: GPU Clusters & HPC

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000