AI HPC Infra Engineer — GPU Clusters & Slurm Expert

The Supreme HR Advisory Pte. Ltd.

Singapore

On-site

SGD 56,000 - 78,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

The Supreme HR Advisory Pte. Ltd. is recruiting an AI Infrastructure Engineer to design and operate high-density multi-GPU compute clusters and optimized AI data paths.

You will architect, deploy, and maintain Linux-based HPC environments with advanced networking, storage, and automation tools, closely collaborating with AI/ML teams. The role requires 3–6+ years of hands-on infra experience, GPU mastery, and proficiency with Kubernetes, Slurm, and IaC toolchains.

Qualifications

  • Bachelor’s degree in Computer Science, IT, Computer Engineering, or equivalent practical experience.
  • 3–6+ years of hands-on experience in infrastructure engineering, HPC, DevOps, or cloud infrastructure.
  • Deep expertise in Linux systems administration, kernel tuning, and shell scripting (Bash/Python).
  • Strong understanding of GPU hardware architectures, CUDA runtimes, and PCIe/NVLink topologies.
  • Hands-on experience with Kubernetes (GPU operator, device plugins) and HPC schedulers (Slurm, Run:ai, Ray).
  • Proven experience with RDMA (RoCE v2 /InfiniBand), PFC, and ECN configurations.
  • Familiarity with high-IOPS, low-latency shared storage architectures for AI datasets and model checkpoints.
  • Automation: Terraform and configuration management (Ansible).
  • Relevant certifications (CKA/CKAD, NVIDIA NC/NP, AWS/Azure/GCP) are a plus.

Responsibilities

  • Architect, configure, and maintain high-density multi-GPU compute clusters (NVIDIA HGX/DGX).
  • Set up container orchestration platforms optimized for AI/ML workloads.
  • Monitor GPU health, telemetry, utilization, and thermals; minimize idle time.
  • Design and optimize low-latency network fabrics for distributed training (InfiniBand, NVLink).
  • Configure and scale high-throughput parallel file systems and object storage for datapipelines.
  • Build and manage automated deployment pipelines using Terraform, Ansible, Helm, or Pulumi.
  • Maintain golden images, Linux tuning, and firmware updates.
  • Implement end-to-end monitoring and dashboards (Prometheus, Grafana, DCGM, NVIDIA SMI).
  • Lead incident response, RCA, and disaster recovery for AI environments.

Skills

Linux OS tuning
GPU hardware architectures
Performance troubleshooting
Distributed training

Education

Bachelor’s Degree in Computer Science/IT/Engineering

Tools

Kubernetes
Slurm
Ray
Terraform
Ansible
Pulumi

Job description

The Supreme HR Advisory Pte. Ltd. is recruiting an AI Infrastructure Engineer to design and operate high-density multi-GPU compute clusters and optimized AI data paths.

You will architect, deploy, and maintain Linux-based HPC environments with advanced networking, storage, and automation tools, closely collaborating with AI/ML teams. The role requires 3–6+ years of hands-on infra experience, GPU mastery, and proficiency with Kubernetes, Slurm, and IaC toolchains.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Infra Architect: GPU Clusters & HPC Ops
Senior AI Infra Architect: GPU Clusters & HPC Ops

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
GPU HPC Infra Engineer for AI Clusters
GPU HPC Infra Engineer for AI Clusters

The Supreme HR Advisory Pte. Ltd. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra Engineer: GPU Clusters & HPC Networking
AI Infra Engineer: GPU Clusters & HPC Networking

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000
AI Systems Infra Engineer - Multi-GPU HPC
AI Systems Infra Engineer - Multi-GPU HPC

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra Engineer: GPU HPC Clusters & Orchestration
AI Infra Engineer: GPU HPC Clusters & Orchestration

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra Architect: GPU Clusters & HPC
AI Infra Architect: GPU Clusters & HPC

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra Engineer: HPC GPU Clusters & Kubernetes
AI Infra Engineer: HPC GPU Clusters & Kubernetes

The Supreme HR Advisory Pte. Ltd. • Singapore

On-site
SGD 56,000 - 78,000
Senior AI Compute & HPC Infrastructure Engineer
Senior AI Compute & HPC Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra DevOps Engineer - GPU/Cloud HPC
AI Infra DevOps Engineer - GPU/Cloud HPC

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 57,000 - 77,000
GPU AI Infra Engineer for High-Perf Clusters
GPU AI Infra Engineer for High-Perf Clusters

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000