Senior HPC Systems Engineer

kaust (king abdullah university of science and technology)

Saudi Arabia

On-site

SAR 360,000 - 600,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

KAUST Supercomputing Laboratory (KSL) seeks a Senior HPC Systems Administrator to oversee a 600-node HPC cluster, high-performance storage, InfiniBand and Ethernet networks, and day-to-day operations supporting researchers across computational science, engineering, big data and AI/ML workloads.

You will install and manage HPC subsystems, deploy Slurm, handle security, performance tuning, and work closely with faculty, application teams, and vendors to ensure reliable, scalable research computing

Qualifications

  • Five years of experience supporting large scale computing platforms and related subsystems.
  • Experience troubleshooting complex hardware issues and documenting root cause analysis.
  • Experience managing parallel storage systems (Lustre, GPFS, Weka, Vast, or similar).
  • Experience benchmarking HPC system components (CPU, memory, InfiniBand, storage).
  • Experience administering workload managers/schedulers (Slurm, LSF, or PBS). Strong Linux system administration experience (RHEL, Rocky Linux, or CentOS).
  • Experience with configuration management tools such as Ansible or Puppet.
  • Ability to coordinate with researchers, application support teams, and vendors to resolve complex issues and drive them to closure.
  • Familiarity with Kubernetes and container orchestration platforms would be desirable

Responsibilities

  • Provide timely user support via multiple channels while maintaining high service standards.
  • Install, configure, and manage HPC subsystems including compute nodes, storage, InfiniBand and Ethernet.
  • Deploy and manage cluster management software, monitoring tools, and supporting services for HPC clusters.
  • Install and administer Slurm, manage QOS policies, accounts, and automation scripts.
  • Develop automation scripts in Bash and Python to streamline tasks.
  • Deploy and manage container environments for HPC workloads.
  • Benchmark HPC components to ensure optimal performance and identify tuning opportunities.
  • Enforce security best practices including node hardening and kernel patching.
  • Manage parallel file systems with performance tuning and capacity planning.
  • Support research activities and collaborate with researchers, application teams and industry partners.
  • Develop software tools and utilities to support research projects on cluster systems.
  • Drive proof-of-concept projects and technology evaluations end-to-end.
  • Coordinate with vendors to resolve issues promptly and maintain documentation.

Skills

Linux system administration
HPC environment
Slurm workload manager
Python
Bash
Ansible
Puppet
Docker
Singularity/Apptainer
Kubernetes
InfiniBand networking
GPU/CPU hardware troubleshooting
Performance benchmarking
Parallel file systems (Lustre/GPFS...)
Documentation & training

Education

Bachelor’s or master’s degree in computer science/engineering, Information Systems, or equivalent

Tools

Ansible
Puppet
Slurm
Docker
Singularity/Apptainer
Kubernetes

Job description

KAUST Supercomputing Laboratory (KSL) seeks a Senior HPC Systems Administrator to oversee a 600-node HPC cluster, high-performance storage, InfiniBand and Ethernet networks, and day-to-day operations supporting researchers across computational science, engineering, big data and AI/ML workloads.

You will install and manage HPC subsystems, deploy Slurm, handle security, performance tuning, and work closely with faculty, application teams, and vendors to ensure reliable, scalable research computing

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Senior Systems Administrator
HPC Senior Systems Administrator

kaust (king abdullah university of science and technology) • Saudi Arabia

On-site
SAR 360,000 - 600,000
Senior HPC Systems Lead & Team Manager
Senior HPC Systems Lead & Team Manager

King Abdullah Bin Abdulaziz University Hospital • Saudi Arabia

On-site
SAR 300,000 - 540,000
HPC Computational Scientist
HPC Computational Scientist

KAUST (King Abdullah University of Science and Technology) • Thuwal

On-site
SAR 700,000 - 1,000,000
HPC Lead Systems Administrator
HPC Lead Systems Administrator

King Abdullah Bin Abdulaziz University Hospital • Saudi Arabia

On-site
SAR 300,000 - 540,000
Senior HPC Computational Scientist — Research & Optimization
Senior HPC Computational Scientist — Research & Optimization

KAUST (King Abdullah University of Science and Technology) • Thuwal

On-site
SAR 700,000 - 1,000,000
Senior Linux System Administrator (HPC)
Senior Linux System Administrator (HPC)

Deepsource Technologies • Riyadh

On-site
SAR 180,000 - 300,000
Senior Platform Lead - Linux Scientific Workstations
Senior Platform Lead - Linux Scientific Workstations

King Abdullah Bin Abdulaziz University Hospital • Saudi Arabia

On-site
SAR 400,000 - 700,000
Research User Computing Lead
Research User Computing Lead

King Abdullah Bin Abdulaziz University Hospital • Saudi Arabia

On-site
SAR 400,000 - 700,000
Senior Linux Systems Engineer - RHEL, Kubernetes, HPC
Senior Linux Systems Engineer - RHEL, Kubernetes, HPC

Deepsource Technologies • Riyadh

On-site
SAR 180,000 - 300,000
Research User Computing Lead
Research User Computing Lead

kaust (king abdullah university of science and technology) • Saudi Arabia

On-site
SAR 350,000 - 520,000