Get more replies from employers
Send a job-specific resume in minutes.
The Supreme HR Advisory Pte Ltd is seeking an AI Infrastructure Engineer to architect and operate high-density GPU compute clusters in a Singapore on-site setting. You will deploy and manage containerized AI/ML workloads using Kubernetes, Slurm, or Ray, while tuning Linux systems and GPUs for peak performance.
You will build automated pipelines with Terraform and Ansible, monitor system health with Prometheus/Grafana, and collaborate with AI/ML teams to minimize bottlenecks during distributed
5 days, Mon - Fri 8.30am to 5.30pm
Salary: $5,000 to $7,000
Location:Kaki Bukit
Architect, configure, and maintain high-density multi-GPU compute clusters (e.g. NVIDIA HGX/DGX architectures).
Implement and manage container orchestration platforms (Kubernetes, Slurm, or Ray) optimized for AI/ML distributed workloads.
Monitor GPU health, telemetry, utilization, and thermals; minimize idle compute time and prevent single-node bottlenecks.
Design and optimize low-latency, lossless network fabrics supporting distributed training (InfiniBand, RoCE v2, NVLink, spine-leaf topologies).
Configure and scale high-throughput parallel file systems and object storage (e.g. Lustre, GPFS/IBM Spectrum Scale, Ceph, MinIO, NVMe-oF) to feed high-speed datapipelines.
Build and manage automated deployment pipelines using Terraform, Ansible, Helm, or Pulumi.
Maintain standard golden images, Linux OS tuning (kernel parameters, NUMA node binding, GPU drivers, CUDA/cuDNN libraries), and firmware updates.
Set up end-to-end monitoring, alerting, and metrics dashboards (Prometheus, Grafana, DCGM exporter, NVIDIA System Management Interface).
Partner with AI/ML engineering teams to diagnose network bottlenecks, NCCL communication latency, and I/O wait states during distributed training jobs.
Lead incident response, root-cause analysis (RCA), and disaster recovery plans for mission-critical AI environments.
Operating Systems: Deep expertise in Linux systems administration, kernel tuning, and shell scripting (Bash/Python).
Accelerated Compute: Strong understanding of GPU hardware architectures, CUDA runtimes, and PCIe/NVLink topologies.
Orchestration & Workload Scheduling: Hands-on experience with Kubernetes (GPU operator, device plugins) and/or HPC schedulers (Slurm, Run:ai, Ray).
High-Speed Networking: Proven experience with RDMA (RoCE v2 /InfiniBand), PFC (Priority Flow Control), and ECN configurations.
Storage Systems: Familiarity with high-IOPS, low-latency shared storage architectures for AI datasets and model checkpoints.
Automation: Proficiency in Infrastructure as Code (Terraform) and configuration management (Ansible).
Bachelor's Degree in Computer Science, Information Technology, Computer
Engineering, or equivalent practical experience.
3-6+ years of hands-on experience in infrastructure engineering, high-performance computing (HPC), DevOps, or cloud infrastructure.
Relevant certifications are a plus (e.g., CKA/CKAD, NVIDIA Certified
Associate/Professional, AWS/Azure/GCP Solutions Architect).