6723 - AI Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD.

Singapore

On-site

SGD 56,000 - 78,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

The Supreme HR Advisory Pte. Ltd. in Singapore is seeking an AI Infrastructure Engineer to architect and maintain high-density GPU compute clusters for AI/ML workloads.

You will optimize Kubernetes, Slurm, and Ray orchestration, while ensuring low-latency networking and fast data access. Your role includes building IaC pipelines with Terraform and Ansible, tuning Linux systems, and collaborating with AI/ML teams to resolve bottlenecks using CUDA/NCCL and HPC tooling.

Qualifications

  • Deep Linux systems administration, kernel tuning, and shell scripting experience.
  • Strong GPU hardware knowledge, CUDA runtimes, and PCIe/NVLink topology awareness.
  • Hands-on Kubernetes (GPU operator) and HPC schedulers (Slurm, Run:AI, Ray).
  • Experience with RDMA networks (RoCE v2/InfiniBand) and ECN/PFC settings.
  • Familiar with shared storage architectures for AI datasets and checkpoints.
  • IaC (Terraform) and configuration management (Ansible) proficiency.
  • Bachelor's degree and 3–6+ years in infrastructure engineering or HPC.

Responsibilities

  • Architect, configure, and maintain high-density multi-GPU compute clusters.
  • Implement and manage container orchestration platforms optimized for AI workloads.
  • Monitor GPU health, utilization, and prevent bottlenecks; optimize I/O.
  • Design low-latency networks and scale parallel file systems for AI data.
  • Build automated deployment pipelines using Terraform, Ansible, and similar tools.
  • Lead incident response and RCA for mission-critical AI environments.

Skills

Linux systems administration
GPU hardware architectures
Kubernetes
Slurm
Ray
Terraform
Ansible
Bash
Python
CUDA / cuDNN
NVIDIA HGX / DGX
NCCL

Education

Bachelor's Degree in Computer Science, Information Technology, Computer Engineering, or equivalent

Tools

Terraform
Ansible
NVIDIA DGX tooling
RDMA tooling

Job description

AI Infrastructure Engineer

5 days, Mon - Fri 8.30am to 5.30pm

Salary: $5,000 to $7,000

Location:Kaki Bukit Ave

Job scopes:
Compute & Cluster Management
  • Architect, configure, and maintain high-density multi-GPU compute clusters (e.g. NVIDIA HGX/DGX architectures).
  • Implement and manage container orchestration platforms (Kubernetes, Slurm, or Ray) optimized for AI/ML distributed workloads.
  • Monitor GPU health, telemetry, utilization, and thermals; minimize idle compute time and prevent single-node bottlenecks.
High-Performance Networking & Storage
  • Design and optimize low-latency, lossless network fabrics supporting distributed training (InfiniBand, RoCE v2, NVLink, spine-leaf topologies).
  • Configure and scale high-throughput parallel file systems and object storage (e.g. Lustre, GPFS/IBM Spectrum Scale, Ceph, MinIO, NVMe-oF) to feed high-speed datapipelines.
Automation & Infrastructure as Code (IaC)
  • Build and manage automated deployment pipelines using Terraform, Ansible, Helm, or Pulumi.
  • Maintain standard golden images, Linux OS tuning (kernel parameters, NUMA node binding, GPU drivers, CUDA/cuDNN libraries), and firmware updates.
Operations, Observability & Performance
  • Set up end-to-end monitoring, alerting, and metrics dashboards (Prometheus, Grafana, DCGM exporter, NVIDIA System Management Interface).
  • Partner with AI/ML engineering teams to diagnose network bottlenecks, NCCL communication latency, and I/O wait states during distributed training jobs.
  • Lead incident response, root-cause analysis (RCA), and disaster recovery plans for mission-critical AI environments.
Requirements:
  • Operating Systems:Deep expertise in Linux systems administration, kernel tuning, and shell scripting (Bash/Python).
  • Accelerated Compute:Strong understanding of GPU hardware architectures, CUDA runtimes, and PCIe/NVLink topologies.
  • Orchestration & Workload Scheduling:Hands-on experience with Kubernetes (GPU operator, device plugins) and/or HPC schedulers (Slurm, Run:ai, Ray).
  • High-Speed Networking:Proven experience with RDMA (RoCE v2 /InfiniBand), PFC (Priority Flow Control), and ECN configurations.
  • Storage Systems:Familiarity with high-IOPS, low-latency shared storage architectures for AI datasets and model checkpoints.
  • Automation:Proficiency in Infrastructure as Code (Terraform) and configuration management (Ansible).
  • Bachelor's Degree in Computer Science, Information Technology, Computer
    Engineering, or equivalent practical experience.
  • 3-6+years of hands-on experience in infrastructure engineering, high-performance computing (HPC), DevOps, or cloud infrastructure.
  • Relevant certifications are a plus (e.g., CKA/CKAD, NVIDIA Certified
    Associate/Professional, AWS/Azure/GCP Solutions Architect).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer - LCYL
AI Infrastructure Engineer - LCYL

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI DevOps Engineer (Cloud Infrastucture)
AI DevOps Engineer (Cloud Infrastucture)

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 57,000 - 77,000
AI Infrastructure Engineer - Kaki Bukit [2683]
AI Infrastructure Engineer - Kaki Bukit [2683]

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer
AI Infrastructure Engineer

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000
AB03 - AI Infrastructure Engineer
AB03 - AI Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer | Up to $7000 | {hkhdv}
AI Infrastructure Engineer | Up to $7000 | {hkhdv}

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
YR23- AI Infrastructure Engineer |HPC/DevOps|CKA/CKAD Cert|Min 3 Years Experience
YR23- AI Infrastructure Engineer |HPC/DevOps|CKA/CKAD Cert|Min 3 Years Experience

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
5-day work week
AI Infrastructure Engineer | Degree | Kaki Bukit | 5 Days | Up To $7K - 4461
AI Infrastructure Engineer | Degree | Kaki Bukit | 5 Days | Up To $7K - 4461

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
6723 - AI Systems Infrastructure Engineer [Up to $7K - Kaki Bukit - Hand On Exp in Infrastructure engineering]
6723 - AI Systems Infrastructure Engineer [Up to $7K - Kaki Bukit - Hand On Exp in Infrastructure engineering]

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer | Up to $7K - 0310
AI Infrastructure Engineer | Up to $7K - 0310

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000