Get more replies from employers
Send a job-specific resume in minutes.
RUNSUN SERVICE PTE. LTD. in Singapore seeks an experienced AI System Engineer to design, deploy, operate, and optimize AI training clusters and GPU platforms. The role involves managing Linux systems, CUDA/NVIDIA drivers, and container orchestration to ensure high availability.
You will build and maintain scalable AI infrastructure, automate operations, and work with datacenter teams to resolve training environment issues, while supporting on-call rotations and occasional travel as required.
We are seeking an experienced AI System Engineer to design, deploy, operate, and optimize AI training clusters, GPU computing platforms, and supporting infrastructure. The ideal candidate should possess strong expertise in Linux systems, GPU computing environments, container platforms, and AI/HPC cluster architectures.
Deploy and operate AI training and HPC clusters;
Install, configure, and optimize operating systems on GPU servers;
Manage cluster resources and capacity;
Perform system upgrades, patch management, and change implementation;
Develop and maintain standardized operational procedures.
Manage large-scale Linux environments;
Perform system performance tuning;
Analyze system logs and kernel issues;
Troubleshoot system stability problems;
Manage user access and security policies.
Manage NVIDIA GPU computing platforms;
Deploy and maintain CUDA, NVIDIA Drivers, and Fabric Manager;
Troubleshoot GPU, NV Link, and NV Switch-related issues;
Optimize GPU cluster performance;
Support customer in resolving training environment issues.
Build and maintain Kubernetes clusters;
Support AI workload scheduling;
Deploy and manage container runtime environments;
Optimize GPU utilization within containers;
Manage Kubernetes high-availability architectures.
Manage distributed storage platforms;
Operate high-speed networking environments;
Collaborate with datacenter teams for troubleshooting;
Monitor infrastructure health and performance;
Improve system reliability and availability.
Develop infrastructure automation tools;
Create deployment and health-check scripts;
Build monitoring and observability platforms;
Implement alerting and self-healing mechanisms;
Improve operational efficiency through automation.
Hands-on experience with Kubernetes;
Familiarity with Docker and Containerd;
Experience with Helm;
Knowledge of GPU Operator;
Understanding of Kubernetes GPU scheduling.
Strong understanding of TCP/IP networking;
Experience with InfiniBand and RoCE;
Familiarity with RDMA architectures;
Experience with one or more storage systems Lustre, BeeGFS, Ceph
and NFS
Strong scripting skills in Shell and Python;
Know about Ansible;
Ability to build infrastructure automation scripts.
Experience supporting Large Language Model(LLM) training platforms;
Knowledge of Slurm workload manager;
Experience with Ray and Kubeflow;
Familiarity with NVIDIA Base Command Manager(BCM);
Experience with NVIDIA NIM;
Knowledge of PXE deployment solutions;
Experience on using DDN product;
Experience operating large-scale GPUclusters;
Experience supporting global datacenteroperations.