Hardware Engineer

RUNSUN SERVICE PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

RUNSUN SERVICE PTE. LTD. seeks an experienced AI Hardware Engineer to support design, deployment, validation, and troubleshooting of AI training clusters, GPU servers, networking, and storage infrastructure.

The ideal candidate will possess strong expertise in server hardware, GPU platforms, high-speed networking, and data center infrastructure to support large-scale AI/HPC environments. The role requires hands-on Linux, scripting, and on-site operational capabilities, with readiness for on-call

Qualifications

  • Bachelor's degree or above in Computer Engineering, Electrical Engineering, Telecommunications, or related fields.
  • Strong knowledge of x86 server architecture.
  • Experience with NVIDIA GPU products (H100, H200, B200, B300).
  • Familiar with CUDA, NCCL, NVLink, NVSwitch and GPU Direct RDMA.
  • Strong Linux administration skills (Ubuntu, Rocky Linux) and scripting (Shell, Python).

Responsibilities

  • Deploy, validate, and maintain AI GPU servers.
  • Analyze system logs, BMC logs, and hardware alerts.
  • Participate in AI/HPC cluster deployment and validation.
  • Configure and maintain high-speed networking and distributed storage.
  • Develop automation scripts for hardware health checks and deployment.
  • Participate in on-call rotation and occasional business trips.

Skills

Linux administration
Python scripting
Shell/Bash scripting
x86 server architecture
NVIDIA GPU platforms
NVIDIA Grace CPU
Networking (TCP/IP, InfiniBand, RoCE)
BMC/IPMI management
On-call readiness

Education

Bachelor's degree in Computer Engineering, Electrical Engineering, Telecommunications, or related fields

Tools

HGX
DGX
GB200 NVL72
GB300 NVL72
CUDA
NCCL
GPU Direct RDMA

Job description

We are seeking an experienced AI Hardware Engineer to support the design, deployment, validation, and troubleshooting of AI training clusters, GPU servers, networking, and storage infrastructure. The ideal candidate should have strong expertise in server hardware, GPU platforms, high-speed networking, and data center infrastructure to support large-scale AI/HPC environments.

Key Responsibilities
AI Server Hardware Management
  • Deploy, validate, and maintain AI GPU servers;
  • Perform hardware diagnostics and component replacement;
  • Analyze system logs, BMC logs, and hardware alerts;
  • Manage server hardware lifecycle.
GPU Platform Support
  • Deploy and validate NVIDIA GPU platforms;
  • Troubleshoot GPU-related;
  • Perform GPU benchmarking and stress testing;
  • Support CUDA, NCCL, and GPU fabric troubleshooting.
AI Cluster Deployment & Validation
  • Participate in AI/HPC cluster deployment;
  • Execute cluster hardware qualification testing;
  • Produce validation reports and documentation.
Network & Storage Support
  • Configure and maintain high-speed networking:
  • Support distributed storage systems:
  • Assist with performance analysis and troubleshooting.
Automation & Tool Development
  • Develop automation scripts for:

    Hardware health checks

    Cluster validation

    Deployment automation

    Log collection

  • Build tools for testing and operations.
  • Good communication, teamwork, and ownership mindset.
  • Willing to participate in on-call rotation, maintenance windows, and emergency incident response, willing to accept short-term business trips.
Required Qualifications

Bachelor's degree or above in Computer Engineering, Electrical Engineering, Telecommunications, or related fields.

Hardware
  • Strong knowledge of x86 server architecture;
  • Familiar with Intel, AMD, and NVIDIA Grace CPU platforms;
  • Experience with:
  • HGX
  • DGX
  • GB200 NVL72
  • GB300 NVL72
  • Knowledge of BMC/IPMI management.
GPU & AI Platform
  • Experience with NVIDIA GPU products H100, H200, B200, B300
  • Familiar with: CUDA ,NCCL ,NV Link ,NV Switch and GPU Direct RDMA
Linux
  • Strong Linux administration skills (Ubuntu, Rocky Linux);
  • Proficient in: Shell ,Python , Bash
  • Capable of independent troubleshooting.
Networking

Strong understanding of: TCP/IP , VLAN , BGP ,OSPF ,RDMA , InfiniBand and RoCE

Preferred Qualities
  • Experience operating AI training clusters; Kubernetes experience; Slurm administration
  • PXE deployment experience; GPU Fabric Manager expertise;
  • Experience with hyperscale AI datacenter deployments.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

System Engineer
System Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Training Cluster Hardware Engineer
AI Training Cluster Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Server Engineer(AI Cluster/GPU)
Server Engineer(AI Cluster/GPU)

PaleBlueDot AI • Singapore

On-site
SGD 120,000 - 180,000
6723 - AI Infrastructure Engineer
6723 - AI Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SwapeTech • Singapore

On-site
SGD 180,000 - 260,000
6723 - HPC Infrastructure Engineer | $5K–$7K | Kaki Bukit | Slurm, GPU & InfiniBand
6723 - HPC Infrastructure Engineer | $5K–$7K | Kaki Bukit | Slurm, GPU & InfiniBand

The Supreme HR Advisory Pte. Ltd. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer - LCYL
AI Infrastructure Engineer - LCYL

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Network Engineer
AI Infrastructure Network Engineer

Tencent • Singapore

On-site
SGD 180,000 - 240,000
Senior AI Infrastructure Support Engineer
Senior AI Infrastructure Support Engineer

nscale operations apac pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000
6723 - GPU Infrastructure Engineer | Up to $7K | Kaki Bukit | NVIDIA, CUDA & HPC
6723 - GPU Infrastructure Engineer | Up to $7K | Kaki Bukit | NVIDIA, CUDA & HPC

The Supreme HR Advisory Pte. Ltd. • Singapore

On-site
SGD 56,000 - 78,000