Senior HPC & AI Cluster Systems Engineer

PaleBlueDot AI

Singapore

On-site

SGD 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

PaleBlueDot AI seeks a seasoned HPC/AI systems engineer to own the full lifecycle of data center servers, from design to validation and ongoing maintenance of HPC/AI clusters.

You will deploy heterogeneous resources (GPUs, XPUs), optimize performance, and lead virtualization/ containerization efforts while building automation with Shell, Python, and Ansible; strong documentation and customer support skills are essential.

Qualifications

  • 5+ years of experience building large HPC clusters with thousands of GPUs.
  • Strong expertise in GPU architecture and parallel computing (MPI, OpenMP).
  • Proficiency in virtualization and containerization; benchmarking and HPC FS (Lustre/GPFS).
  • HPC certifications are preferred and English proficiency is required.

Responsibilities

  • Manage the full lifecycle of data center servers, including design, deployment, performance tuning and validation for HPC/AI clusters.
  • Lead deployment of heterogeneous computing resources (GPUs, XPUs) and optimize system performance and stability.
  • Prepare technical documentation, provide customer support, and drive automation with Shell, Python and Ansible.

Skills

HPC clusters
GPU architecture
MPI
OpenMP
Shell scripting
Python
Ansible
virtualization
containerization
NCCL benchmarking
Lustre/GPFS

Tools

MPI
OpenMP
NCCL
Lustre
GPFS

Job description

PaleBlueDot AI seeks a seasoned HPC/AI systems engineer to own the full lifecycle of data center servers, from design to validation and ongoing maintenance of HPC/AI clusters.

You will deploy heterogeneous resources (GPUs, XPUs), optimize performance, and lead virtualization/ containerization efforts while building automation with Shell, Python, and Ansible; strong documentation and customer support skills are essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Server Engineer(AI Cluster)
Server Engineer(AI Cluster)

PaleBlueDot AI • Singapore

On-site
SGD 120,000 - 170,000
Senior HPC Systems Engineer - Linux Clusters & AI Workloads
Senior HPC Systems Engineer - Linux Clusters & AI Workloads

OPENSOURCE TECHNOLOGIES PTE. LTD. • Singapore

On-site
SGD 90,000 - 140,000
Senior AI/HPC Systems Engineer
Senior AI/HPC Systems Engineer

NVIDIA Gruppe • Singapore

On-site
SGD 120,000 - 180,000
Senior Solutions Architect (AI, HPC & Cloud)
Senior Solutions Architect (AI, HPC & Cloud)

Gateway Search • Singapore

On-site
SGD 120,000 - 180,000
System Engineer
System Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Hardware Engineer
Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Infrastructure Engineer - GPU HPC & Kubernetes Expert
AI Infrastructure Engineer - GPU HPC & Kubernetes Expert

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra Architect: GPU Clusters & HPC
AI Infra Architect: GPU Clusters & HPC

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
HPC Systems & Cluster Infra Engineer
HPC Systems & Cluster Infra Engineer

D L RESOURCES PTE LTD • Singapore

On-site
SGD 60,000 - 96,000
HPC High Performance Computing IT Infra Engineer
HPC High Performance Computing IT Infra Engineer

D L RESOURCES PTE LTD • Singapore

On-site
SGD 60,000 - 96,000