Get more replies from employers
Send a job-specific resume in minutes.
D L RESOURCES PTE LTD is seeking an HPC Systems Administrator to support day-to-day operations of on-premise clusters and cloud-backed HPC workloads. You will monitor performance, manage job scheduling, and optimize resources across Linux environments, storage and networks.
Responsibilities include installation, patching, incident response, and collaboration with senior engineers on scaling and tuning HPC infrastructure. Prior HPC exposure is preferred.
Client: Research & Education Sector
Support day-to-day operations of HPC clusters, including compute nodes, storage systems, and high-speed networking infrastructure
Monitor system performance, workload execution, job scheduling, and resource utilization across distributed environments
* Perform installation, configuration, patching, and maintenance of Linux/Unix operating systems (RHEL, CentOS, Ubuntu)
* Administer physical and virtualized server environments (x86 architecture, VMware/KVM)
* Support workload management systems and job schedulers (Slurm, PBS, LSF) for batch processing and resource allocation
* Conduct system health checks, log analysis, troubleshooting, and incident management (root cause analysis)
* Manage user accounts, access control, and authentication systems (LDAP, Active Directory integration)
* Assist in provisioning and configuration of HPC environments, including software stack deployment and environment modules
* Collaborate with senior engineers on cluster optimization, scaling, and performance tuning initiatives
* Maintain system documentation, operational procedures, and runbooks
1–3 years of experience in system administration, infrastructure support, or IT operations
Hands-on experience with Linux/Unix systems administration and command-line environments
* Basic exposure to HPC, distributed systems, or parallel computing environments
* Strong understanding of networking fundamentals (TCP/IP, DNS, SSH, firewall concepts)
* Knowledge of storage technologies (NAS, SAN, distributed/parallel file systems – basic awareness)
* Scripting and automation using Bash/Shell and/or Python
* Familiarity with monitoring and logging tools (e.g., Nagios, Zabbix, Prometheus, Grafana, ELK Stack)
* Strong troubleshooting, analytical thinking, and incident resolution skills
Exposure to GPU computing environments (NVIDIA GPUs, CUDA – basic awareness)
Familiarity with cloud platforms and HPC workloads on cloud (AWS, Azure)
* Awareness of container technologies (Docker) and basic DevOps practices
The project environment is primarily on-premise, with approximately 80% of the infrastructure hosted on Dell and servers, and approximately 20% involving AWS-based HPC/GPU server environments. This provides exposure to both traditional data centre infrastructure and cloud-based GPU/HPC workloads.
of: