HPC High Performance Computing IT Infra Engineer

D L RESOURCES PTE LTD

Singapore

On-site

SGD 60,000 - 96,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

D L RESOURCES PTE LTD is seeking an HPC Systems Administrator to support day-to-day operations of on-premise clusters and cloud-backed HPC workloads. You will monitor performance, manage job scheduling, and optimize resources across Linux environments, storage and networks.

Responsibilities include installation, patching, incident response, and collaboration with senior engineers on scaling and tuning HPC infrastructure. Prior HPC exposure is preferred.

Qualifications

  • 1–3 years of experience in system administration, infrastructure support, or IT operations.
  • Hands-on experience with Linux/Unix systems administration and command-line environments.
  • Basic exposure to HPC, distributed systems, or parallel computing environments.
  • Strong understanding of networking fundamentals (TCP/IP, DNS, SSH, firewall concepts).
  • Knowledge of storage technologies (NAS, SAN, distributed/parallel file systems – basic awareness).
  • Scripting and automation using Bash/Shell and/or Python.
  • Familiarity with monitoring and logging tools (e.g., Nagios, Zabbix, Prometheus, Grafana, ELK Stack).
  • Strong troubleshooting, analytical thinking, and incident resolution skills.

Responsibilities

  • Support day-to-day operations of HPC clusters, including compute nodes, storage systems, and high-speed networking infrastructure.
  • Monitor system performance, workload execution, job scheduling, and resource utilization across distributed environments.
  • Perform installation, configuration, patching, and maintenance of Linux/Unix operating systems (RHEL, CentOS, Ubuntu).
  • Administer physical and virtualized server environments (x86 architecture, VMware/KVM).
  • Support workload management systems and job schedulers (Slurm, PBS, LSF) for batch processing and resource allocation.
  • Conduct system health checks, log analysis, troubleshooting, and incident management (root cause analysis).
  • Manage user accounts, access control, and authentication systems (LDAP, Active Directory integration).
  • Assist in provisioning and configuration of HPC environments, including software stack deployment and environment modules.
  • Collaborate with senior engineers on cluster optimization, scaling, and performance tuning initiatives.
  • Maintain system documentation, operational procedures, and runbooks

Skills

Linux/Unix Administration
Bash/Python Scripting
HPC / Parallel Computing
Networking (TCP/IP, DNS, SSH)
Monitoring & Logging (Nagios, Zabbix,-
Batch Scheduling (Slurm, PBS, LSF)
VMware/KVM virtualization
Storage (NAS/SAN/GPFS/Lustre)
Docker/Kubernetes basics
Incident management

Tools

VMware
KVM
Docker
Kubernetes
Splunk
ELK Stack

Job description

Client: Research & Education Sector

Key Responsibilities

Support day-to-day operations of HPC clusters, including compute nodes, storage systems, and high-speed networking infrastructure
Monitor system performance, workload execution, job scheduling, and resource utilization across distributed environments
* Perform installation, configuration, patching, and maintenance of Linux/Unix operating systems (RHEL, CentOS, Ubuntu)
* Administer physical and virtualized server environments (x86 architecture, VMware/KVM)
* Support workload management systems and job schedulers (Slurm, PBS, LSF) for batch processing and resource allocation
* Conduct system health checks, log analysis, troubleshooting, and incident management (root cause analysis)
* Manage user accounts, access control, and authentication systems (LDAP, Active Directory integration)
* Assist in provisioning and configuration of HPC environments, including software stack deployment and environment modules
* Collaborate with senior engineers on cluster optimization, scaling, and performance tuning initiatives
* Maintain system documentation, operational procedures, and runbooks

Required Skills & Experience

1–3 years of experience in system administration, infrastructure support, or IT operations
Hands-on experience with Linux/Unix systems administration and command-line environments
* Basic exposure to HPC, distributed systems, or parallel computing environments
* Strong understanding of networking fundamentals (TCP/IP, DNS, SSH, firewall concepts)
* Knowledge of storage technologies (NAS, SAN, distributed/parallel file systems – basic awareness)
* Scripting and automation using Bash/Shell and/or Python
* Familiarity with monitoring and logging tools (e.g., Nagios, Zabbix, Prometheus, Grafana, ELK Stack)
* Strong troubleshooting, analytical thinking, and incident resolution skills

Preferred Skills

Exposure to GPU computing environments (NVIDIA GPUs, CUDA – basic awareness)
Familiarity with cloud platforms and HPC workloads on cloud (AWS, Azure)
* Awareness of container technologies (Docker) and basic DevOps practices

Project / Environment Tech Stack. (Mostly On-Prem)

The project environment is primarily on-premise, with approximately 80% of the infrastructure hosted on Dell and servers, and approximately 20% involving AWS-based HPC/GPU server environments. This provides exposure to both traditional data centre infrastructure and cloud-based GPU/HPC workloads.

of:

80% On-Premise HPC Infrastructure
  • Dell servers, Huawei servers, physical data centre infrastructure
  • x86 server architecture, CPU-based compute infrastructure, bare-metal servers
  • Virtualized server environments using VMware / KVM where applicable
  • Linux/Unix-based systems including RHEL, CentOS, Ubuntu, AIX
  • HPC cluster components across compute, storage, and networking layers
  • High-performance storage / parallel file systems such as IBM Spectrum Scale (GPFS), Lustre, BeeGFS; with NAS / SAN / NFS exposure where applicable
  • Cluster networking and connectivity including TCP/IP, DNS, SSH, firewall concepts, high-throughput Ethernet; InfiniBand / RDMA exposure preferred
  • HPC workload scheduling and batch processing using Slurm, PBS, or LSF
  • Monitoring and logging tools such as Nagios, Zabbix, Prometheus, Grafana, ELK Stack, or Splunk
20% AWS HPC / GPU-Related Environment
  • AWS cloud infrastructure supporting HPC / GPU-related workloads
  • AWS GPU server exposure, including NVIDIA GPU-based compute instances
  • GPU computing environment exposure, including NVIDIA CUDA, NVIDIA drivers, GPU utilization, and GPU workload monitoring where applicable
  • Cloud HPC workload support involving compute, storage, networking, and security configurations
  • Exposure to AWS HPC services or related tools such as AWS ParallelCluster, AWS Batch, EC2 GPU instances, EBS / FSx / S3 storage, VPC, IAM, and security groups
  • Container and DevOps exposure where applicable, including Docker, Kubernetes, and basic CI/CD or automation practices
  • Hybrid HPC environment exposure involving on-premise infrastructure integrated with AWS-based compute or GPU resources
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC(High Performance Computing) Engineer
HPC(High Performance Computing) Engineer

Infinite Computer Solutions Pte Ltd • Singapore

On-site
SGD 70,000 - 100,000
Systems Engineer (HPC)
Systems Engineer (HPC)

OPENSOURCE TECHNOLOGIES PTE. LTD. • Singapore

On-site
SGD 90,000 - 140,000
System ENgineer (HPC)
System ENgineer (HPC)

OPENSOURCE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Server Engineer(AI Cluster)
Server Engineer(AI Cluster)

PaleBlueDot AI • Singapore

On-site
SGD 120,000 - 170,000
System Engineer
System Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Systems Engineer (HPC)
Systems Engineer (HPC)

FUJITSU ASIA PTE LTD • Singapore

On-site
SGD 90,000 - 150,000
Hardware Engineer
Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior Infrastructure Engineer – IDC, Bare-Metal Kubernetes
Senior Infrastructure Engineer – IDC, Bare-Metal Kubernetes

Jobtailor • Singapore

On-site
SGD 120,000 - 190,000
HPC Systems & Cluster Infra Engineer
HPC Systems & Cluster Infra Engineer

D L RESOURCES PTE LTD • Singapore

On-site
SGD 60,000 - 96,000
AI Infrastructure Engineer
AI Infrastructure Engineer

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000