Server Engineer(AI Cluster)

PaleBlueDot AI

Singapore

On-site

SGD 120,000 - 170,000

Full time

12 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

PaleBlueDot AI seeks a seasoned HPC/AI systems engineer to own the full lifecycle of data center servers, from design to validation and ongoing maintenance of HPC/AI clusters.

You will deploy heterogeneous resources (GPUs, XPUs), optimize performance, and lead virtualization/ containerization efforts while building automation with Shell, Python, and Ansible; strong documentation and customer support skills are essential.

Qualifications

  • 5+ years of experience building large HPC clusters with thousands of GPUs.
  • Strong expertise in GPU architecture and parallel computing (MPI, OpenMP).
  • Proficiency in virtualization and containerization; benchmarking and HPC FS (Lustre/GPFS).
  • HPC certifications are preferred and English proficiency is required.

Responsibilities

  • Manage the full lifecycle of data center servers, including design, deployment, performance tuning and validation for HPC/AI clusters.
  • Lead deployment of heterogeneous computing resources (GPUs, XPUs) and optimize system performance and stability.
  • Prepare technical documentation, provide customer support, and drive automation with Shell, Python and Ansible.

Skills

HPC clusters
GPU architecture
MPI
OpenMP
Shell scripting
Python
Ansible
virtualization
containerization
NCCL benchmarking
Lustre/GPFS

Tools

MPI
OpenMP
NCCL
Lustre
GPFS

Job description

  • Manage the full lifecycle of data center servers, including design, deployment, performance tuning, and validation, as well as the implementation, operation, and maintenance of HPC/AI clusters.
  • Lead the deployment of heterogeneous computing resources such as GPUs and XPUs, optimize system performance and stability, and monitor and maintain server operations.
  • Prepare technical documentation, provide customer support, and drive the development of automation scripts using Shell, Python, and Ansible.
Key Requirements
  • At least 5 years of relevant experience, with hands‑on expertise in building large‑scale HPC clusters with thousands of GPUs, GPU hardware architecture, and parallel computing technologies such as MPI and OpenMP.
  • Strong proficiency in virtualization and containerization technologies, with experience in HPL and NCCL benchmarking and high‑performance file systems such as Lustre and GPFS.
  • HPC‑related certifications are preferred.
  • Professional working proficiency in English.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC & AI Cluster Systems Engineer
Senior HPC & AI Cluster Systems Engineer

PaleBlueDot AI • Singapore

On-site
SGD 120,000 - 170,000
Hardware Engineer
Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
System Engineer
System Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Hardware Engineer
Hardware Engineer

runsun cloud pte ltd • Singapore

On-site
SGD 100,000 - 160,000
System ENgineer (HPC)
System ENgineer (HPC)

OPENSOURCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Systems Engineer (HPC)
Systems Engineer (HPC)

OPENSOURCE TECHNOLOGIES PTE. LTD. • Singapore

On-site
SGD 90,000 - 140,000
Senior HPC Systems Engineer - Linux Clusters & AI Workloads
Senior HPC Systems Engineer - Linux Clusters & AI Workloads

OPENSOURCE TECHNOLOGIES PTE. LTD. • Singapore

On-site
SGD 90,000 - 140,000
HPC AI Lead Engineer
HPC AI Lead Engineer

National University Of Singapore • Singapore

On-site
SGD 80,000 - 120,000
System Engineer
System Engineer

runsun cloud pte ltd • Singapore

On-site
SGD 120,000 - 180,000
Systems Engineer (HPC)
Systems Engineer (HPC)

FUJITSU ASIA PTE LTD • Singapore

On-site
SGD 90,000 - 150,000