Contract NVIDIA H800 Support Engineer (Location: Cyberport)

OneAsia Network Limited

Hong Kong

On-site

HKD 600,000 - 900,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

OneAsia Network Limited in Hong Kong is seeking a skilled NVIDIA H800 Support Engineer to manage enterprise-scale AI supercomputing environments. You will provide deep technical troubleshooting for DGX SuperPOD deployments across multi-node clusters.

You bring strong Linux system administration, high-performance networking and experience with NVIDIA GPU tools, InfiniBand, Slurm or Kubernetes, and Python/Bash scripting to automate operations.

Qualifications

  • Bachelor's or Master's degree in computer science, computer engineering, or other engineering fields or equivalent experience.
  • 3+ years in providing in-depth system support and debugging experience.
  • System level expertise of CPU/GPU server architecture, NICs, Linux, system software and kernel drivers.
  • Proven hands-on experience in Linux troubleshooting with good problem identification, resolution and solving skills.
  • Experience with NVIDIA GPU infrastructure tools (DCGM, nvidia-smi, MIG, NVSwitch diagnostics)
  • Hands-on experience with InfiniBand protocols, RDMA, and RoCEv2 architectures.
  • Experience with cluster orchestrators like Slurm or Kubernetes.
  • Proficiency in Bash and Python scripting for system monitoring and automation.

Responsibilities

  • Diagnose and resolve multi-node hardware and software failures across clusters.
  • Guide customer discussions on network design, compute/storage and support bring up of server/network/cluster deployments.
  • Troubleshoot and fine-tune ultra-high-bandwidth fabrics to maximize compute efficiency.
  • Perform deep-dive post-mortems on systemic bugs, node drops and cluster regressions.
  • Partner with vendor support teams to patch software bugs and systemvulnerability.

Skills

Linux troubleshooting
NVIDIA DGX
NVIDIA GPU tools
InfiniBand RDMA
Slurm Kubernetes
Bash scripting
Python scripting
Cluster management
Networking design

Education

Bachelor/Master in CS/CE

Tools

DCGM
nvidia-smi
NCCL tuning
Docker
Kubernetes

Job description

OneAsia is a leading IT services and solution provider in Asia providing cloud based solution as well as data centre services. OneAsia's top-tier rated data centres are located across Asia to keep customers connected from anywhere in the world with consistent levels of quality, security and service. Partnering with the technology leaders, OneAsia is able to offer a full range of cloud computing solutions, from infrastructure, management to application software to business of all sizes without additional capital investment or strong IT support. Flexibility, reliability, and security are the core values of OneAsia. With fully redundant infrastructure, well developed systems, multi-layered security and skilled personnel, OneAsia delivers professional and reliable services to customers.

Role Overview

TheNVIDIAH800SupportEngineer manages enterprise-scale AI supercomputing environments. You will provide deep technical troubleshooting for NVIDIA DGX SuperPOD deployments. This role bridges advanced system administration, high-performance networking, and distributed AI software stacks.

Key Responsibilities

Diagnose and resolve multi-node hardware and software failures across clusters.

Guide customer discussions on network design, compute/storage and support bring up of server/network/cluster deployments.

Troubleshoot and fine-tune ultra-high-bandwidth fabrics to maximize compute efficiency.

Perform deep-dive post-mortems on systemic bugs, node drops and cluster regressions.

Partner with vendor support teams to patch software bugs and systemvulnerability.

Required Qualifications

Bachelor's or Master's degree in computer science, Computer Engineering, or other Engineering fields or equivalent experience.

3+ years in providing in-depth System support and debugging experience, strong organizational skills and ability to prioritize/multi-task easily with limited supervision.

System level expertise of CPU/GPU server architecture, NICs, Linux, system software and kernel drivers

Proven hands-on experience in Linux troubleshooting with good problem identification, resolution and solving skills.

Experience with NVIDIA GPU infrastructure tools (DCGM, nvidia-smi, MIG, NVSwitch diagnostics)

Hands-on experience with InfiniBand protocols, RDMA, and RoCEv2 architectures.

Experience with cluster orchestrators like Slurm or Kubernetes.

Proficiency in Bash and Python scripting for system monitoring and automation.

Preferred Qualifications

Hands-on experience with NVIDIA Base Command Manager, Slurm orchestrators, and tuning NVIDIA NCCL for distributed deep learning.

Experience with container technologies (Docker, Apptainer/Singularity, Kubernetes).

Exposure to high-throughput parallel filesystems like GPFS or Lustre.

Attractive remuneration package and fringe benefit will be offered to the right candidate.

For more information about us, please visit the corporate website athttps://www.oneas1a.com/

Equal employment opportunities apply to all applicants. All applications and data collected will be treated in strict confidence and used exclusively for recruitment purposes. Only short-listed candidates will be invited for interview. The company will retain the applications for a maximum period of 12 months and may refer suitable candidates to other vacancies within the Group.

Your application will include the following questions:

  • Which of the following statements best describes your right to work in Hong Kong?
  • What's your expected monthly basic salary?

OneAsia is a leading IT services and solution provider in Asia providing cloud based solution as well as data centre services. OneAsia's top-tier rated data centres are located across Asia to keep our customers connected from anywhere in the world with consistent

levels of quality, security and service.

Partnering with the technology leaders, OneAsia is able to offer a full range of cloud computing solutions, from infrastructure, management to application software to business of all sizes without additional capital investment or strong IT support. Furthermore,

OneAsia can customize data centre services such as colocation, managed services, optimization and business continuity based on customer requirements.

Flexibility, reliability, and security are the core values of OneAsia. With fully redundant infrastructure, well developed systems, multi-layered security and skilled personnel, OneAsia delivers professional and reliable services to customers. With an aim to

keep customers connected wherever and whenever they are, OneAsia is staying at the forefront of the industry with extensive infrastructure coverage in Greater China, Singapore, Malaysia and Vietnam.

OneAsia Network Limited

OneAsia is a leading IT services and solution provider in Asia providing cloud based solution as well as data centre services. OneAsia's top-tier rated data centres are located across Asia to keep our customers connected from anywhere in the world with consistent

levels of quality, security and service.

Partnering with the technology leaders, OneAsia is able to offer a full range of cloud computing solutions, from infrastructure, management to application software to business of all sizes without additional capital investment or strong IT support. Furthermore,

OneAsia can customize data centre services such as colocation, managed services, optimization and business continuity based on customer requirements.

Flexibility, reliability, and security are the core values of OneAsia. With fully redundant infrastructure, well developed systems, multi-layered security and skilled personnel, OneAsia delivers professional and reliable services to customers. With an aim to

keep customers connected wherever and whenever they are, OneAsia is staying at the forefront of the industry with extensive infrastructure coverage in Greater China, Singapore, Malaysia and Vietnam.

OneAsia Network Limited

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Contract Infrastructure Security Engineer (Location: Cyberport)
Contract Infrastructure Security Engineer (Location: Cyberport)

OneAsia Network Limited • Hong Kong

On-site
HKD 600,000 - 900,000
NVIDIA H800 AI Supercomputing Support Engineer
NVIDIA H800 AI Supercomputing Support Engineer

OneAsia Network Limited • Hong Kong

On-site
HKD 600,000 - 900,000
Senior Network Deployment Engineer - APAC
Senior Network Deployment Engineer - APAC

NVIDIA • Hong Kong

On-site
HKD 900,000 - 1,300,000
Cloud Engineer (Huawei Product, North Point, over $40K)
Cloud Engineer (Huawei Product, North Point, over $40K)

CL Technical Services Ltd. • Hong Kong

On-site
HKD 320,000 - 520,000
IT Support / Proximity Engineer (Contract)
IT Support / Proximity Engineer (Contract)

ServiceOne Limited • Hong Kong

On-site
HKD 167,000 - 290,000
Banking holidays
Medical insurance
Senior Manager, Supercomputing Center
Senior Manager, Supercomputing Center

HSITP Hong Kong-Shenzhen Innovation and Technology Park • Hong Kong

On-site
HKD 1,000,000 - 1,400,000
Competitive annual leave entitlement
Medical benefits from Day-1 with depen
Training sponsorship
+1
Associate Director, Supercomputing Center (Technical & Product)
Associate Director, Supercomputing Center (Technical & Product)

HSITP Hong Kong-Shenzhen Innovation and Technology Park • Hong Kong

On-site
HKD 1,200,000 - 2,400,000
Competitive annual leave
Medical benefits from Day-1
Training sponsorship
+1
Server Engineer (HK - Hybrid)
Server Engineer (HK - Hybrid)

Pragmatike • Hong Kong

Hybrid
HKD 480,000 - 720,000
Network Engineer
Network Engineer

Eclipse Trading • Hong Kong

On-site
HKD 700,000 - 900,000
Dynamic environment
Flat structure
Work-life balance
+2
Server Engineer (On-site Hong Kong)
Server Engineer (On-site Hong Kong)

Pragmatike • Hong Kong

On-site
HKD 420,000 - 660,000