HPC Senior Systems Administrator

KAUST (King Abdullah University of Science and Technology)

Makkah Region

On-site

SAR 260,000 - 420,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

KAUST (King Abdullah University of Science and Technology) in Saudi Arabia is seeking a Senior HPC Systems Administrator to manage a ~600-node HPC cluster, storage, and networks, supporting researchers across computational science, engineering, big data, and AI/ML workloads.

You will install, configure, and administer Slurm, Lustre/GPFS storage, container runtimes (Singularity/Apptainer, Docker), and security hardening, while collaborating with researchers, vendors, and application support teams

Qualifications

  • Bachelor’s or master’s degree in computer science/engineering, information systems, or equivalent.
  • Five years of experience supporting large-scale computing platforms.
  • Experience administering workload managers (Slurm, LSF, or PBS).
  • Experience with parallel file systems (Lustre, GPFS, Weka, Vast).
  • Strong Linux system administration experience (RHEL, Rocky Linux, or CentOS).
  • Experience with configuration management tools such as Ansible or Puppet.
  • Ability to coordinate with researchers, application support teams, and vendors to resolve complex issues.

Responsibilities

  • Provide user support via telephone, walk-in, email, and ticketing system while maintaining high service standards.
  • Install, configure, and manage HPC subsystems including compute nodes, storage, InfiniBand, Ethernet, and CM tools.
  • Deploy and manage cluster management software, monitoring tools, and supporting services for HPC clusters.
  • Install and administer Slurm, manage QOS policies, accounts, and automation scripts (Python and C++).
  • Develop and maintain automation scripts in Bash and Python for admin tasks.
  • Deploy and manage container environments (Singularity/Apptainer, Docker) for HPC workloads.
  • Benchmark HPC components to ensure optimal performance and identify tuning opportunities.

Skills

Linux system administration
HPC systems & cluster management
Python scripting
Bash scripting
MPI/OpenMP/CUDA programming models
English communication

Education

Bachelor’s or Master’s degree in Computer Science/Engineering/Information Systems

Tools

Slurm workload manager
Ansible
Puppet
Docker
Singularity/Apptainer
InfiniBand networks
GPFS/Lustre/Vast
Kubernetes

Job description

We are seeking a highly motivated and skilled Senior HPC Systems Administrator to join the KAUST Supercomputing Laboratory (KSL). The successful candidate will be responsible for managing an HPC cluster of approximately 600 CPU and GPU nodes, HPC storage systems, InfiniBand and Ethernet networks, and day-to-day operational issues. The role provides broad support to researchers and end-users across computational science, engineering, big data analysis, and artificial intelligence/machine learning workloads.

Major Responsibilities – include but are not limited to -
  • Provide timely and effective user support via telephone, walk-in, email, and ticketing system for all inquiry types while maintain high customer service standards.
  • Install, configure, and manage HPC subsystems including compute nodes, high-performance storage systems, InfiniBand, Ethernet, and configuration management tools (e.g., Ansible, Puppet).
  • Deploy and manage cluster management software, monitoring tools, and supporting services for operating HPC clusters.
  • Install and administer the Slurm workload manager, manage QOS policies, accounts, accounting, and related automation scripts (Python and C++).
  • Develop and maintain automation scripts in Bash and Python to streamline system administration tasks.
  • Deploy and manage container environments (Singularity/Apptainer, Docker) for HPC workloads.
  • Benchmark HPC system components like CPU, memory, InfiniBand, and storage periodically to ensure optimal performance and identify tuning opportunities across hardware, driver, and application layers.
  • Enforce security best practices including node hardening, kernel patching, and compliance across all systems.
  • Manage parallel file systems such as Lustre, GPFS, Weka, or Vast, including performance tuning and capacity planning.
  • Directly support research activities in computational science, engineering, data analysis, and AI/ML by working closely with faculty, researchers, collaboration partners, and industrial partners in collaboration with application support teams.
  • Develop software tools and utilities as needed to support research projects on cluster systems and subsystems.
  • Drive proof-of-concept projects and technology evaluations end-to-end and research industry best practices and advocate system enhancements.
  • Coordinate with vendors and third-party service providers to report and resolve issues in a timely manner.
  • Develop and maintain user documentation, standard operating procedures, and training materials in the internal wiki.
  • Stay at the forefront of HPC advancements through continuous learning, industry conferences, and professional collaboration, while driving benchmarking initiatives to inform future hardware procurement.
  • Expertise in supporting users of computational science and engineering, data analysis, and artificial intelligence applications and libraries in different HPC environments.
  • Strong expertise in Linux system administration (RHEL, Rocky Linux, or CentOS) in large-scale HPC environments.
  • Proficiency with HPC applications and programming models (Fortran, C/C++, Python, MPI, OpenMP, CUDA, OpenACC).
  • Demonstrated track record of managing complex HPC systems, including parallel file systems, job schedulers, InfiniBand/Ethernet networks, and monitoring systems.
  • Experience with configuration management tools (Ansible, Puppet, or equivalent).
  • Familiarity with computational science, data analysis, and AI/ML applications and libraries used in HPC environments.
  • Knowledge of project management principles and practices.
  • Demonstrated ability to support research activities in a highly collaborative HPC environment.
  • Strong analytical, problem-solving, and decision-making skills.
  • Proactively identifies and implements system improvements; takes initiative and sees tasks through to closure.
  • Ability to manage multiple concurrent projects and deliver high-quality results within deadlines.
  • Proven ability to collaborate cross-functionally with researchers, application teams, and vendors.
  • Effective in multi-cultural, international work environments.
  • Excellent verbal and written communication skills in English, including the ability to prepare and deliver technical reports and presentations.
Qualifications

Bachelor’s or master’s degree in computer science/engineering, Information Systems, or equivalent

Experience
Experience requirement, minimum:
  • Five years of experience supporting large scale computing platforms and related subsystems.
  • Experience troubleshooting complex hardware issues and documenting root cause analysis.
  • Experience managing parallel storage systems (Lustre, GPFS, Weka, Vast, or similar).
  • Experience benchmarking HPC system components (CPU, memory, InfiniBand, storage).
  • Experience administering workload managers/schedulers (Slurm, LSF, or PBS).
  • Strong Linux system administration experience (RHEL, Rocky Linux, or CentOS).
  • Experience with configuration management tools such as Ansible or Puppet.
  • Ability to coordinate with researchers, application support teams, and vendors to resolve complex issues and drive them to closure.
  • Proven ability to collaborate cross-functionally and drive initiatives to completion.
  • Familiarity with Kubernetes and container orchestration platforms would be desirable
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Computational Scientist
HPC Computational Scientist

KAUST (King Abdullah University of Science and Technology) • Thuwal

On-site
SAR 700,000 - 1,000,000
HPC Lead Systems Administrator at King Abdullah University of Science and Technology
HPC Lead Systems Administrator at King Abdullah University of Science and Technology

King Abdullah University of Science and Technology • Saudi Arabia

On-site
SAR 450,000 - 750,000
HPC Lead Systems Administrator
HPC Lead Systems Administrator

King Abdullah Bin Abdulaziz University Hospital • Saudi Arabia

On-site
SAR 300,000 - 540,000
HPC Lead System Administrator
HPC Lead System Administrator

KAUST (King Abdullah University of Science and Technology) • Thuwal

On-site
SAR 300,000 - 450,000
Senior HPC Systems Administrator for Research Clusters
Senior HPC Systems Administrator for Research Clusters

KAUST (King Abdullah University of Science and Technology) • Makkah Region

On-site
SAR 260,000 - 420,000
AI-ML Support Analyst
AI-ML Support Analyst

KAUST (King Abdullah University of Science and Technology) • Makkah Region

On-site
SAR 180,000 - 300,000
AI/ML Automation Analyst
AI/ML Automation Analyst

KAUST (King Abdullah University of Science and Technology) • Makkah Region

On-site
SAR 180,000 - 240,000
Senior Linux System Administrator (HPC)- Riyadh, KSA
Senior Linux System Administrator (HPC)- Riyadh, KSA

DS DeepSource • Riyadh

On-site
SAR 180,000 - 240,000
Research User Computing Lead
Research User Computing Lead

King Abdullah Bin Abdulaziz University Hospital • Saudi Arabia

On-site
SAR 400,000 - 700,000
HPC Data Center Hardware Specialist (Cray/GPU Ops)
HPC Data Center Hardware Specialist (Cray/GPU Ops)

KAUST (King Abdullah University of Science and Technology) • Makkah Region

On-site
SAR 1,200,000 - 1,800,000