6723 - AI Systems Infrastructure Engineer [Up to $7K - Kaki Bukit - Hand On Exp in Infrastructure engineering]

THE SUPREME HR ADVISORY PTE. LTD.

Singapore

On-site

SGD 56,000 - 78,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

The Supreme HR Advisory Pte. Ltd. is seeking an AI Infrastructure Engineer to join our Singapore team at Kaki Bukit. This role focuses on designing and operating high-density multi-GPU compute clusters, with emphasis on container orchestration for AI/ML workloads.

You will optimize infrastructure, networking, and storage for distributed training, monitor system health, and collaborate with AI/ML engineers. Requires 3–6+ years in infra engineering and relevant certifications are a plus.

Qualifications

  • Bachelor's degree or equivalent practical experience in CS/IT/Engineering.
  • 3–6+ years hands-on infra engineering, HPC, DevOps, or cloud infra.
  • Relevant certifications (CKA/CKAD, NVIDIA/AWS/Azure/GCP) are a plus.

Responsibilities

  • Architect, configure, and maintain high-density GPU compute clusters (NVIDIA HGX/DGX).
  • Deploy and manage container orchestration platforms (Kubernetes, Slurm, Ray) for AI/ML workloads.
  • Monitor GPU health and system telemetry; optimize performance and prevent bottlenecks.
  • Design and implement IaC and automation (Terraform, Ansible).

Skills

Linux admin
Kernel tuning
Shell scripting
Bash
Python

Education

Bachelor's Degree in Computer Science/IT/Engineering

Tools

Kubernetes
Slurm
Ray

Job description

AI Infrastructure Engineer

5 days, Mon - Fri 8.30am to 5.30pm

Salary: $5,000 to $7,000

Location: Kaki Bukit

Operating Systems

Deep expertise in Linux systems administration, kernel tuning, and shell scripting (Bash/Python).

Accelerated Compute

Strong understanding of GPU hardware architectures, CUDA runtimes, and PCIe/NVLink topologies.

Orchestration & Workload Scheduling

Hands-on experience with Kubernetes (GPU operator, device plugins) and/or HPC schedulers (Slurm, Run:ai, Ray).

High-Speed Networking

Proven experience with RDMA (RoCE v2 /InfiniBand), PFC (Priority Flow Control), and ECN configurations.

Storage Systems

Familiarity with high-IOPS, low-latency shared storage architectures for AI datasets and model checkpoints.

Automation

Proficiency in Infrastructure as Code (Terraform) and configuration management (Ansible).

Qualifications
  • Bachelor's Degree in Computer Science, Information Technology, Computer Engineering, or equivalent practical experience.
  • 3-6+ years of hands-on experience in infrastructure engineering, high-performance computing (HPC), DevOps, or cloud infrastructure.
  • Relevant certifications are a plus (e.g., CKA/CKAD, NVIDIA Certified Associate/Professional, AWS/Azure/GCP Solutions Architect).
Job scopes
Compute & Cluster Management
  • Architect, configure, and maintain high-density multi-GPU compute clusters (e.g. NVIDIA HGX/DGX architectures).
  • Implement and manage container orchestration platforms (Kubernetes, Slurm, or Ray) optimized for AI/ML distributed workloads.
  • Monitor GPU health, telemetry, utilization, and thermals; minimize idle compute time and prevent single-node bottlenecks.
High-Performance Networking & Storage
  • Design and optimize low-latency, lossless network fabrics supporting distributed training (InfiniBand, RoCE v2, NVLink, spine-leaf topologies).
  • Configure and scale high-throughput parallel file systems and object storage (e.g. Lustre, GPFS/IBM Spectrum Scale, Ceph, MinIO, NVMe-oF) to feed high-speed data pipelines.
Automation & Infrastructure as Code (IaC)
  • Build and manage automated deployment pipelines using Terraform, Ansible, Helm, or Pulumi.
  • Maintain standard golden images, Linux OS tuning (kernel parameters, NUMA node binding, GPU drivers, CUDA/cuDNN libraries), and firmware updates.
Operations, Observability & Performance
  • Set up end-to-end monitoring, alerting, and metrics dashboards (Prometheus, Grafana, DCGM exporter, NVIDIA System Management Interface).
  • Partner with AI/ML engineering teams to diagnose network bottlenecks, NCCL communication latency, and I/O wait states during distributed training jobs.
  • Lead incident response, root-cause analysis (RCA), and disaster recovery plans for mission-critical AI environments.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Systems Infrastructure Engineer - Up to $7K - 0310
AI Systems Infrastructure Engineer - Up to $7K - 0310

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer | Up to $7K - 0310
AI Infrastructure Engineer | Up to $7K - 0310

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
6723 - AI Systems Engineer | Up to $7K | Kaki Bukit | Kubernetes, Linux & GPU Clusters
6723 - AI Systems Engineer | Up to $7K | Kaki Bukit | Kubernetes, Linux & GPU Clusters

The Supreme HR Advisory Pte. Ltd. • Singapore

On-site
SGD 56,000 - 78,000
6723 - AI Infrastructure Engineer
6723 - AI Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
6723 - GPU Infrastructure Engineer | Up to $7K | Kaki Bukit | NVIDIA, CUDA & HPC
6723 - GPU Infrastructure Engineer | Up to $7K | Kaki Bukit | NVIDIA, CUDA & HPC

The Supreme HR Advisory Pte. Ltd. • Singapore

On-site
SGD 56,000 - 78,000
6723 - HPC Infrastructure Engineer | $5K–$7K | Kaki Bukit | Slurm, GPU & InfiniBand
6723 - HPC Infrastructure Engineer | $5K–$7K | Kaki Bukit | Slurm, GPU & InfiniBand

The Supreme HR Advisory Pte. Ltd. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer | Up to $7000 | {hkhdv}
AI Infrastructure Engineer | Up to $7000 | {hkhdv}

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer - LCYL
AI Infrastructure Engineer - LCYL

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI GPU Infra Engineer — Scale High-Performance Compute
AI GPU Infra Engineer — Scale High-Performance Compute

Hamilton Barnes Associates Limited • Singapore

On-site
SGD 180,000 - 220,000
Comprehensive health, dental, and vision insurance
401(k) with employer matching
Equity options
+1
AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SwapeTech • Singapore

On-site
SGD 180,000 - 260,000