AI Infrastructure Engineer - LCYL

THE SUPREME HR ADVISORY PTE. LTD.

Singapore

On-site

SGD 56,000 - 78,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

THE SUPREME HR ADVISORY PTE. LTD. seeks an AI Infrastructure Engineer to design and operate high-density GPU compute clusters and distributed AI workloads. You will implement GPUs, storage, and networking with modern orchestrators across Linux environments.

You will optimize performance, reliability, and scalability, collaborating with ML teams. Possess strong Linux, CUDA, and HPC tooling, with familiar Terraform/Ansible for IaC and config management.

Qualifications

  • Linux system administration, kernel tuning, and Bash/Python scripting.
  • Deep understanding of GPU hardware, CUDA runtimes, PCIe/NVLink topologies.
  • Kubernetes with GPU operator and HPC schedulers (Slurm, Run:ai, Ray).
  • RDMA (RoCE v2/InfiniBand), PFC, ECN configurations.

Responsibilities

  • Architect, configure, and maintain high-density GPU compute clusters.
  • Manage container orchestration platforms (Kubernetes, Slurm, Ray) for AI workloads.
  • Monitor GPU health, telemetry, utilization, and thermals; minimize idle compute time.

Skills

Linux systems administration
GPU architectures
CUDA runtimes
Kubernetes (GPU operator)
HPC schedulers (Slurm, Run:AI, Ray)
RDMA/InfiniBand networking
Terraform
Ansible
Shell scripting (Bash/Python)
Storage design (Lustre/Ceph)

Education

Bachelor’s Degree in Computer Science/IT/Engineering or equivalent

Tools

Kubernetes
Slurm
Ray
Terraform
Ansible
Helm
Pulumi
Lustre
Ceph
NVIDIA DCGM

Job description

AI Infrastructure Engineer

5 days, Mon - Fri 8.30am to 5.30pm

Salary: $5,000 to $7,000

Location:2 Kaki Bukit Ave 1, Singapore 417938

Job scopes:
Compute & Cluster Management
  • Architect, configure, and maintain high-density multi-GPU compute clusters (e.g. NVIDIA HGX/DGX architectures).
  • Implement and manage container orchestration platforms (Kubernetes, Slurm, or Ray) optimized for AI/ML distributed workloads.
  • Monitor GPU health, telemetry, utilization, and thermals; minimize idle compute time and prevent single-node bottlenecks.
High-Performance Networking & Storage
  • Design and optimize low-latency, lossless network fabrics supporting distributed training (InfiniBand, RoCE v2, NVLink, spine-leaf topologies).
  • Configure and scale high-throughput parallel file systems and object storage (e.g. Lustre, GPFS/IBM Spectrum Scale, Ceph, MinIO, NVMe-oF) to feed high-speed datapipelines.
Automation & Infrastructure as Code (IaC)
  • Build and manage automated deployment pipelines using Terraform, Ansible, Helm, or Pulumi.
  • Maintain standard golden images, Linux OS tuning (kernel parameters, NUMA node binding, GPU drivers, CUDA/cuDNN libraries), and firmware updates.
Operations, Observability & Performance
  • Set up end-to-end monitoring, alerting, and metrics dashboards (Prometheus, Grafana, DCGM exporter, NVIDIA System Management Interface).
  • Partner with AI/ML engineering teams to diagnose network bottlenecks, NCCL communication latency, and I/O wait states during distributed training jobs.
  • Lead incident response, root-cause analysis (RCA), and disaster recovery plans for mission-critical AI environments.
Requirements:
  • Operating Systems: Deep expertise in Linux systems administration, kernel tuning, and shell scripting (Bash/Python).
  • Accelerated Compute: Strong understanding of GPU hardware architectures, CUDA runtimes, and PCIe/NVLink topologies.
  • Orchestration & Workload Scheduling: Hands-on experience with Kubernetes (GPU operator, device plugins) and/or HPC schedulers (Slurm, Run:ai, Ray).
  • High-Speed Networking: Proven experience with RDMA (RoCE v2 /InfiniBand), PFC (Priority Flow Control), and ECN configurations.
  • Storage Systems: Familiarity with high-IOPS, low-latency shared storage architectures for AI datasets and model checkpoints.
  • Automation: Proficiency in Infrastructure as Code (Terraform) and configuration management (Ansible).
  • Bachelor’s Degree in Computer Science, Information Technology, Computer
    Engineering, or equivalent practical experience.
  • 3–6+ years of hands-on experience in infrastructure engineering, high-performance computing (HPC), DevOps, or cloud infrastructure.
  • Relevant certifications are a plus (e.g., CKA/CKAD, NVIDIA Certified Associate/Professional, AWS/Azure/GCP Solutions Architect).
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer
AI Infrastructure Engineer

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000
YR23- AI Infrastructure Engineer |HPC/DevOps|CKA/CKAD Cert|Min 3 Years Experience
YR23- AI Infrastructure Engineer |HPC/DevOps|CKA/CKAD Cert|Min 3 Years Experience

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
5-day work week
AB03 - AI Infrastructure Engineer
AB03 - AI Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer | Degree | Kaki Bukit | 5 Days | Up To $7K - 4461
AI Infrastructure Engineer | Degree | Kaki Bukit | 5 Days | Up To $7K - 4461

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer | Up to $7K - 0310
AI Infrastructure Engineer | Up to $7K - 0310

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SwapeTech • Singapore

On-site
SGD 180,000 - 260,000
Senior AI Compute & HPC Infrastructure Engineer
Senior AI Compute & HPC Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
Hardware Engineer
Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Infrastructure Project Engineer
AI Infrastructure Project Engineer

RN CARE PTE. LTD. • Singapore

On-site
SGD 50,000 - 70,000
GPU HPC Infrastructure Engineer
GPU HPC Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
5-day work week