Hardware Engineer

runsun cloud pte ltd

Singapore

On-site

SGD 100,000 - 160,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

runsun cloud pte ltd seeks an experienced AI Hardware Engineer to design, deploy, validate, and troubleshoot AI training clusters, GPU servers, networking, and storage infrastructure in a Singapore data center. You will work on server hardware, GPU platforms, and high-speed networking to support large-scale AI/HPC environments.

The role involves on-site deployment, diagnostics, and automation, with on-call rotation and potential short business trips.

Qualifications

  • Bachelor's degree or higher in Computer/Electrical/Telecommunications or related fields.
  • Strong Linux administration and independent troubleshooting capability.
  • Experience with AI/HPC clusters and GPU platforms is preferred.

Responsibilities

  • Deploy, validate, and maintain AI GPU servers.
  • Perform hardware diagnostics and component replacement.
  • Analyze system logs, BMC logs, and hardware alerts.
  • Manage server hardware lifecycle.
  • Deploy and validate NVIDIA GPU platforms; troubleshoot GPU-related issues; perform benchmarking and stress tests.
  • Configure high-speed networking and support storage for AI clusters.
  • Develop automation scripts for health checks, deployment, and log collection.

Skills

Linux administration
Troubleshooting
Scripting

Education

Bachelor's degree or above in Computer/Electrical/Telecommunications or related fields

Tools

HGX/DGX platforms
BMC/IPMI
NVIDIA GPU platforms
CUDA/NCCL

Job description

Job Description & Requirements

We are seeking an experienced AI Hardware Engineer to support the design, deployment, validation, and troubleshooting of AI training clusters, GPU servers, networking, and storage infrastructure. The ideal candidate should have strong expertise in server hardware, GPU platforms, high-speed networking, and data center infrastructure to support large-scale AI/HPC environments.

Key Responsibilities
AI Server Hardware Management
  • Deploy, validate, and maintain AI GPU servers;
  • Perform hardware diagnostics and component replacement;
  • Analyze system logs, BMC logs, and hardware alerts;
  • Manage server hardware lifecycle.
GPU Platform Support
  • Deploy and validate NVIDIA GPU platforms;
  • Troubleshoot GPU-related;
  • Perform GPU benchmarking and stress testing;
  • Support CUDA, NCCL, and GPU fabric troubleshooting.
AI Cluster Deployment & Validation
  • Participate in AI/HPC cluster deployment;
  • Execute cluster hardware qualification testing;
  • Produce validation reports and documentation.
Network & Storage Support
  • Configure and maintain high-speed networking:
  • Support distributed storage systems:
  • Assist with performance analysis and troubleshooting.
Automation & Tool Development
  • Develop automation scripts for:
  • Hardware health checks:
  • Cluster validation
  • Deployment automation
  • Log collection
  • Build tools for testing and operations.
  • Good communication, teamwork, and ownership mindset.
  • Willing to participate in on-call rotation, maintenance windows, and emergency incident response, willing to accept short-term business trips.
Required Qualifications

Bachelor's degree or above in Computer Engineering, Electrical Engineering, Telecommunications, or related fields.

Hardware
  • Strong knowledge of x86 server architecture;
  • Familiar with Intel, AMD, and NVIDIA Grace CPU platforms;
  • Experience with HGX, DGX,GB200NVL72 and GB300 NVL72
  • Knowledge of BMC/IPMI management.
GPU & AI Platform
  • Experience with NVIDIA GPU products H100, H200, B200, B300
  • Familiar with: CUDA ,NCCL ,NV Link ,NV Switch and GPU Direct RDMA
Linux
  • Strong Linux administration skills (Ubuntu, Rocky Linux);
  • Proficient in: Shell ,Python , Bash
  • Capable of independent troubleshooting.
Networking:

Strong understanding of: TCP/IP , VLAN , BGP ,OSPF ,RDMA , InfiniBand and RoCE

Preferred Qualities
  • Experience operating AI training clusters; Kubernetes experience; Slurm administration
  • PXE deployment experience; GPU Fabric Manager expertise;
  • Experience with hyperscale AI datacenter deployments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hardware Engineer
Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
System Engineer
System Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Network Engineer (Data Centre / GPU Infrastructure)
Network Engineer (Data Centre / GPU Infrastructure)

Visa Hunt • Singapore

On-site
SGD 90,000 - 140,000
System Engineer
System Engineer

runsun cloud pte ltd • Singapore

On-site
SGD 120,000 - 180,000
AI Training Cluster Hardware Engineer
AI Training Cluster Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Network Engineer (AI DC)
Network Engineer (AI DC)

Vouch Recruitment • Singapore

On-site
SGD 120,000 - 160,000
AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SwapeTech • Singapore

On-site
SGD 180,000 - 260,000
Server Engineer(AI Cluster)
Server Engineer(AI Cluster)

PaleBlueDot AI • Singapore

On-site
SGD 120,000 - 170,000
Staff Engineer (Hardware Design)
Staff Engineer (Hardware Design)

Sanmina • Singapore

On-site
SGD 70,000 - 100,000
AI Infrastructure Project Engineer
AI Infrastructure Project Engineer

RN CARE PTE. LTD. • Singapore

On-site
SGD 50,000 - 70,000