System Engineer

runsun cloud pte ltd

Singapore

On-site

SGD 120,000 - 180,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

RUNSUN CLOUD PTE LTD is seeking an experienced AI System Engineer to design, deploy, operate, and optimize AI training clusters, GPU computing platforms, and supporting infrastructure in Singapore. The role emphasizes Linux expertise, GPU environments, containers, and HPC architectures.

Responsibilities include deploying AI training clusters, managing NVIDIA GPUs, CUDA drivers, and Fabric Manager, plus Kubernetes orchestration, automation, and platform development.

Qualifications

  • Bachelor's degree or above in Computer Engineering, Electrical Engineering, Telecommunications, or related fields.
  • 3+ years of Linux administration experience; strong knowledge of Ubuntu, Rocky Linux, and RHEL; troubleshooting skills.
  • Familiarity with CUDA, NCCL, NV Link, NV Switch, GPU Direct RDMA; understanding of distributed AI training architectures.
  • Experience with NVIDIA GPU products H100, H200, B200, B300 and related tools (Fabric Manager, drivers).
  • Hands-on Kubernetes, Docker, Helm; GPU scheduling and container runtimes; scripting in Shell and Python.

Responsibilities

  • Deploy and operate AI training and HPC clusters.
  • Install, configure, and optimize operating systems on GPU servers.
  • Manage cluster resources and capacity.
  • Perform system upgrades, patch management, and change implementation.
  • Develop and maintain standardized operational procedures.
  • Manage large-scale Linux environments and perform system performance tuning.
  • Analyze system logs and kernel issues; troubleshoot stability problems.
  • Manage user access and security policies.
  • Manage NVIDIA GPU platforms; deploy CUDA, drivers, and Fabric Manager; troubleshoot GPU issues.
  • Build and maintain Kubernetes clusters; support AI workload scheduling; ensure high availability.
  • Monitor infrastructure health and performance; automate operations; improve efficiency.

Skills

Linux systems
GPU computing
Kubernetes
NVIDIA GPUs
CUDA/NVLink
Shell/Python
Docker
Networking
Storage
Automation tooling

Education

Bachelor's degree or above in Computer/Electrical/Telecommunications or related field

Tools

Kubernetes
Docker
Helm
NVIDIA CUDA drivers
NVIDIA Base Command Manager

Job description

Job Description & Requirements

We are seeking an experienced AI System Engineer to design, deploy, operate, and optimize AI training clusters, GPU computing platforms, and supporting infrastructure. The ideal candidate should possess strong expertise in Linux systems, GPU computing environments, container platforms, and AI/HPC cluster architectures.


Key Responsibilities

AI Cluster deployment and operation


  • Deploy and operate AI training and HPC clusters;

  • Install, configure, and optimize operating systems on GPU servers;

  • Manage cluster resources and capacity;

  • Perform system upgrades, patch management, and change implementation;

  • Develop and maintain standardized operational procedures.


Linux System Management


  • Manage large-scale Linux environments;

  • Perform system performance tuning;

  • Analyze system logs and kernel issues;

  • Troubleshoot system stability problems;

  • Manage user access and security policies.


GPU Platform Support


  • Manage NVIDIA GPU computing platforms;

  • Deploy and maintain CUDA, NVIDIA Drivers, and Fabric Manager;

  • Troubleshoot GPU, NV Link, and NV Switch-related issues;

  • Optimize GPU cluster performance;

  • Support customer in resolving training environment issues.


Container Platform & Orchestration System


  • Build and maintain Kubernetes clusters;

  • Support AI workload scheduling;

  • Deploy and manage container runtime environments;

  • Optimize GPU utilization within containers;

  • Manage Kubernetes high-availability architectures.


AI infrastructure management


  • Manage distributed storage platforms;

  • Operate high-speed networking environments;

  • Collaborate with datacenter teams for troubleshooting;

  • Monitor infrastructure health and performance;

  • Improve system reliability and availability.


Automated operation and maintenance and platform development


  • Develop infrastructure automation tools;

  • Create deployment and health-check scripts;

  • Build monitoring and observability platforms;

  • Implement alerting and self-healing mechanisms;

  • Improve operational efficiency through automation.



  • Good communication, teamwork, and ownership mindset.

  • Willing to participate in on-call rotation, maintenance windows, and emergency incident response, willing to accept short-term business trips.


Required Qualifications


  • Bachelor's degree or above in Computer Engineering, Electrical Engineering, Telecommunications, or related fields.


  • Linux System
    3+years of Linux administration experience;
    Strong knowledge of Ubuntu, Rocky Linux, and RHEL;
    Familiarity with system boot process, kernel, filesystems, and performance tuning;
    Ability to troubleshoot complex system issues independently.




  • GPU & AI Platform
    Strong understanding of NVIDIA GPU architecture;
    Experience with NVIDIA GPU products H100,H200, B200 B300 ,GB200 NVL72 and GB300 NVL72




  • Familiar with CUDA, NCCL, NV Link, NV Switch, GPU Direct RDMA


  • Understanding of distributed AI training architectures.



Container and Cloud Native


  • Hands-on experience with Kubernetes;

  • Familiarity with Docker and Container;

  • Experience with Helm;

  • Knowledge of GPU Operator;

  • Understanding of Kubernetes GPU scheduling.


Networking and Storage


  • Strong understanding of TCP/IP networking;

  • Experience with InfiniBand and RoCE;

  • Familiarity with RDMA architectures;

  • Experience with one or more storage systems Lustre, BeeGFS, Ceph and NFS


Automation capabilities


  • Strong scripting skills in Shell and Python;

  • Know about Ansible;

  • Ability to build infrastructure automation scripts.


Preferred Qualities


  • Experience supporting Large Language Model(LLM) training platforms;

  • Knowledge of Slurm workload manager;

  • Experience with Ray and Kubeflow;

  • Familiarity with NVIDIA Base Command Manager(BCM);

  • Experience with NVIDIA NIM;

  • Knowledge of PXE deployment solutions;

  • Experience on using DDN product;

  • Experience operating large-scale GPU clusters;

  • Experience supporting global datacenter operations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

System Engineer
System Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Hardware Engineer
Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Engineer (ML Systems & Infrastructure)
AI Engineer (ML Systems & Infrastructure)

SwapeTech • Singapore

On-site
SGD 180,000 - 260,000
Hardware Engineer
Hardware Engineer

runsun cloud pte ltd • Singapore

On-site
SGD 100,000 - 160,000
Network Operations Engineer
Network Operations Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
TECHNICAL MANAGER- GPU CLOUD & AI INFRASTRUCTURE
TECHNICAL MANAGER- GPU CLOUD & AI INFRASTRUCTURE

zy future international pte. ltd. • Singapore

On-site
SGD 167,000 - 223,000
HPC Engineer
HPC Engineer

NVIDIA • Singapore

On-site
SGD 120,000 - 180,000
Server Engineer(AI Cluster)
Server Engineer(AI Cluster)

PaleBlueDot AI • Singapore

On-site
SGD 120,000 - 170,000
AI DevOps Engineer
AI DevOps Engineer

GTS Consulting • Singapore

On-site
SGD 100,000 - 180,000
HPC Engineer
HPC Engineer

NVIDIA Gruppe • Singapore

On-site
SGD 120,000 - 180,000