NVIDIA H800 AI Supercomputing Support Engineer

OneAsia Network Limited

Hong Kong

On-site

HKD 600,000 - 900,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

OneAsia Network Limited in Hong Kong is seeking a skilled NVIDIA H800 Support Engineer to manage enterprise-scale AI supercomputing environments. You will provide deep technical troubleshooting for DGX SuperPOD deployments across multi-node clusters.

You bring strong Linux system administration, high-performance networking and experience with NVIDIA GPU tools, InfiniBand, Slurm or Kubernetes, and Python/Bash scripting to automate operations.

Qualifications

  • Bachelor's or Master's degree in computer science, computer engineering, or other engineering fields or equivalent experience.
  • 3+ years in providing in-depth system support and debugging experience.
  • System level expertise of CPU/GPU server architecture, NICs, Linux, system software and kernel drivers.
  • Proven hands-on experience in Linux troubleshooting with good problem identification, resolution and solving skills.
  • Experience with NVIDIA GPU infrastructure tools (DCGM, nvidia-smi, MIG, NVSwitch diagnostics)
  • Hands-on experience with InfiniBand protocols, RDMA, and RoCEv2 architectures.
  • Experience with cluster orchestrators like Slurm or Kubernetes.
  • Proficiency in Bash and Python scripting for system monitoring and automation.

Responsibilities

  • Diagnose and resolve multi-node hardware and software failures across clusters.
  • Guide customer discussions on network design, compute/storage and support bring up of server/network/cluster deployments.
  • Troubleshoot and fine-tune ultra-high-bandwidth fabrics to maximize compute efficiency.
  • Perform deep-dive post-mortems on systemic bugs, node drops and cluster regressions.
  • Partner with vendor support teams to patch software bugs and systemvulnerability.

Skills

Linux troubleshooting
NVIDIA DGX
NVIDIA GPU tools
InfiniBand RDMA
Slurm Kubernetes
Bash scripting
Python scripting
Cluster management
Networking design

Education

Bachelor/Master in CS/CE

Tools

DCGM
nvidia-smi
NCCL tuning
Docker
Kubernetes

Job description

OneAsia Network Limited in Hong Kong is seeking a skilled NVIDIA H800 Support Engineer to manage enterprise-scale AI supercomputing environments. You will provide deep technical troubleshooting for DGX SuperPOD deployments across multi-node clusters.

You bring strong Linux system administration, high-performance networking and experience with NVIDIA GPU tools, InfiniBand, Slurm or Kubernetes, and Python/Bash scripting to automate operations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Contract NVIDIA H800 Support Engineer (Location: Cyberport)
Contract NVIDIA H800 Support Engineer (Location: Cyberport)

OneAsia Network Limited • Hong Kong

On-site
HKD 600,000 - 900,000
Solution Manager - Global Tech (HPC & AI Infrastructure | 5-8m Bonus | Perm)
Solution Manager - Global Tech (HPC & AI Infrastructure | 5-8m Bonus | Perm)

ADECCO Personnel Limited • Hong Kong

On-site
HKD 600,000 - 1,000,000
Contract Infra Security Engineer: Linux & Kubernetes GPU
Contract Infra Security Engineer: Linux & Kubernetes GPU

OneAsia Network Limited • Hong Kong

On-site
HKD 600,000 - 900,000
Senior Network Deployment Engineer - APAC
Senior Network Deployment Engineer - APAC

NVIDIA Corporation • Hong Kong Island

On-site
HKD 900,000 - 1,500,000
AI Platform Engineer: Scalable AI Infra & Scheduler
AI Platform Engineer: Scalable AI Infra & Scheduler

Lenovo • Hong Kong

On-site
HKD 350,000 - 550,000
Contactable email via Lenovo
AI System Engineer (Cyberport site)
AI System Engineer (Cyberport site)

OneAsia Network Limited • Hong Kong

On-site
HKD 600,000 - 900,000
Senior Network Deployment Engineer - APAC
Senior Network Deployment Engineer - APAC

NVIDIA • Hong Kong

On-site
HKD 900,000 - 1,300,000
AI Infra Platform Engineer: Scale Kubernetes & AI Jobs
AI Infra Platform Engineer: Scale Kubernetes & AI Jobs

LPS • Hong Kong

Hybrid
HKD 420,000 - 600,000
AI & HPC Solutions Architect — Global Infrastructure
AI & HPC Solutions Architect — Global Infrastructure

ADECCO Personnel Limited • Hong Kong

On-site
HKD 600,000 - 1,000,000
AI Systems Infrastructure Engineer
AI Systems Infrastructure Engineer

OneAsia Network Limited • Hong Kong

On-site
HKD 600,000 - 900,000