AI Infrastructure Engineer | Up to $7000 | {hkhdv}

THE SUPREME HR ADVISORY PTE. LTD.

Singapore

On-site

SGD 56,000 - 78,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

THE SUPREME HR ADVISORY PTE. LTD. is seeking an AI Infrastructure Engineer to architect and operate high-density GPU clusters and distributed AI workloads. You will implement container orchestration (Kubernetes, Slurm, Ray), optimize kernel parameters, and ensure low-latency networking and fast storage for training pipelines.

You will collaborate with AI/ML teams, monitor GPU health and NCCL latency, and drive automation via Terraform and Ansible to enable reliable, scalable AI infrastructure.

Qualifications

  • Bachelor's degree in CS/IT/CE or equivalent.
  • 3–6+ years hands-on experience in infrastructure engineering, HPC, DevOps, or cloud infra.
  • Deep Linux administration, kernel tuning, and scripting (Bash/Python).
  • Strong understanding of GPU hardware architectures, CUDA runtimes, and PCIe/NVLink topologies.
  • Hands-on experience with Kubernetes GPU operators and HPC schedulers (Slurm, Run:ai, Ray).
  • Proven experience with RDMA (RoCE v2 /InfiniBand), PFC, ECN.
  • Familiarity with high-IOPS, low-latency shared storage for AI datasets

Responsibilities

  • Architect, configure, and maintain AI compute clusters.
  • Implement and manage container orchestration platforms optimized for AI/ML workloads.
  • Monitor GPU health, telemetry, utilization and thermals; minimize idle compute time.
  • Design and optimize low-latency networks and scalable storage for AI pipelines.
  • Build automated deployment pipelines using Terraform and Ansible; maintain golden images.
  • Lead incident response, RCA, and disaster recovery for AI environments.

Skills

Linux administration
Kernel tuning
Shell scripting
GPU architectures
CUDA runtimes
PCIe/NVLink
Kubernetes
Slurm
Run:AI / Ray
NVIDIA HGX/DGX
RDMA / InfiniBand
Storage for AI data
Terraform
Ansible
Monitoring tooling

Education

Bachelor’s Degree in Computer Science / IT / Computer Engineering

Tools

Kubernetes GPU Operator
Slurm HPC Scheduler
Ray
Terraform
Ansible
Pulumi
nvidia-smi / CUDA toolkit

Job description

AI Infrastructure Engineer

5 days, Mon - Fri 8.30am to 5.30pm

Salary: $5,000 to $7,000

Location:Kaki Bukit

Job scopes:
  • Architect, configure, and maintain high-density multi-GPU compute clusters (e.g. NVIDIA HGX/DGX architectures).
  • Implement and manage container orchestration platforms (Kubernetes, Slurm, or Ray) optimized for AI/ML distributed workloads.
  • Monitor GPU health, telemetry, utilization, and thermals; minimize idle compute time and prevent single-node bottlenecks.
High-Performance Networking & Storage
  • Design and optimize low-latency, lossless network fabrics supporting distributed training (InfiniBand, RoCE v2, NVLink, spine-leaf topologies).
  • Configure and scale high-throughput parallel file systems and object storage (e.g. Lustre, GPFS/IBM Spectrum Scale, Ceph, MinIO, NVMe-oF) to feed high-speed datapipelines.
Automation & Infrastructure as Code (IaC)
  • Build and manage automated deployment pipelines using Terraform, Ansible, Helm, or Pulumi.
  • Maintain standard golden images, Linux OS tuning (kernel parameters, NUMA node binding, GPU drivers, CUDA/cuDNN libraries), and firmware updates.
Operations, Observability & Performance
  • Set up end-to-end monitoring, alerting, and metrics dashboards (Prometheus, Grafana, DCGM exporter, NVIDIA System Management Interface).
  • Partner with AI/ML engineering teams to diagnose network bottlenecks, NCCL communication latency, and I/O wait states during distributed training jobs.
  • Lead incident response, root-cause analysis (RCA), and disaster recovery plans for mission-critical AI environments.
Requirements:
  • Operating Systems:Deep expertise in Linux systems administration, kernel tuning, and shell scripting (Bash/Python).
  • Accelerated Compute:Strong understanding of GPU hardware architectures, CUDA runtimes, and PCIe/NVLink topologies.
  • Orchestration & Workload Scheduling:Hands-on experience with Kubernetes (GPU operator, device plugins) and/or HPC schedulers (Slurm, Run:ai, Ray).
  • High-Speed Networking:Proven experience with RDMA (RoCE v2 /InfiniBand), PFC (Priority Flow Control), and ECN configurations.
  • Storage Systems:Familiarity with high-IOPS, low-latency shared storage architectures for AI datasets and model checkpoints.
  • Automation:Proficiency in Infrastructure as Code (Terraform) and configuration management (Ansible).
  • Bachelor’s Degree in Computer Science, Information Technology, Computer
    Engineering, or equivalent practical experience.
  • 3–6+years of hands-on experience in infrastructure engineering, high-performance computing (HPC), DevOps, or cloud infrastructure.
  • Relevant certifications are a plus (e.g., CKA/CKAD, NVIDIA Certified
    Associate/Professional, AWS/Azure/GCP Solutions Architect).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

6723 - AI Infrastructure Engineer
6723 - AI Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer | Degree | Kaki Bukit | 5 Days | Up To $7K - 4461
AI Infrastructure Engineer | Degree | Kaki Bukit | 5 Days | Up To $7K - 4461

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer - LCYL
AI Infrastructure Engineer - LCYL

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer | Up to $7K - 0310
AI Infrastructure Engineer | Up to $7K - 0310

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI DevOps Engineer (Cloud Infrastucture)
AI DevOps Engineer (Cloud Infrastucture)

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 57,000 - 77,000
AI Systems Infrastructure Engineer | Up to $7K
AI Systems Infrastructure Engineer | Up to $7K

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer - Kaki Bukit [2683]
AI Infrastructure Engineer - Kaki Bukit [2683]

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer
AI Infrastructure Engineer

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000
6723 - AI Systems Infrastructure Engineer [Up to $7K - Kaki Bukit - Hand On Exp in Infrastructure engineering]
6723 - AI Systems Infrastructure Engineer [Up to $7K - Kaki Bukit - Hand On Exp in Infrastructure engineering]

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Systems Infrastructure Engineer - Up to $7K - 0310
AI Systems Infrastructure Engineer - Up to $7K - 0310

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000