Senior HPC Engineer – IFM

The Chronicle Of Higher Education, Inc.

United Arab Emirates

On-site

AED 223,200 - 334,800

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

The Chronicle Of Higher Education, Inc. is seeking a Senior HPCEngineer for MBZUAI’s Institute of Foundation Models. This role involves providing technical leadership in designing and operating large-scale GPU infrastructure for AI research.

Responsibilities include optimizing GPU clusters, troubleshooting systems, mentoring junior engineers, and collaborating with researchers. Ideal candidates should have relevant experience in HPC and a degree in a related field.

Qualifications

  • 5+ years in HPC, Linux infrastructure, or large-scale production environments.
  • Experience with Slurm and Linux administration.
  • Experience troubleshooting compute, storage, and networking systems.

Responsibilities

  • Lead operation and optimization of large-scale GPU clusters.
  • Drive reliability, scalability, and performance improvements.
  • Mentor junior engineers.
  • Participate in major incident management and escalations.

Skills

HPC expertise
Linux administration
Cloud infrastructure
Distributed systems
Troubleshooting skills

Education

Bachelor’s degree in relevant field
Master’s degree preferred

Tools

Slurm
NVIDIA technologies
InfiniBand networking
Cloud platforms (AWS, Azure, GCP)
Terraform
Ansible

Job description

Senior HPCEngineer

MBZUAI’s Institute of Foundation Models is seeking a Senior HPCEngineer to provide technical leadership in designing, operating, and evolving large‑scale GPU infrastructure supporting frontier AI research. The Institute for Foundation Models (IFM) operates one of the world’s largest AI‑focused supercomputing environments and is looking for an experienced HPC engineer to contribute to groundbreaking research and development.

Key Responsibilities
  • Lead operation and optimization of large‑scale GPU clusters.
  • Drive reliability, scalability, and performance improvements.
  • Lead troubleshooting and root cause analysis of complex issues.
  • Design and validate new cluster deployments and upgrades.
  • Collaborate with researchers to optimize distributed AI training.
  • Lead vendor engagement and technical reviews.
  • Mentor junior engineers.
  • Define monitoring, operational standards, and capacity planning processes.
  • Participate in major incident management and escalations.
Academic Qualification
  • Bachelor’s degree in computer science, computer engineering, electrical engineering, software engineering, information technology, applied mathematics, physics, or related disciplines.
  • Master’s degree preferred.
Professional Experience Required
  • 5+ years in HPC, Linux infrastructure, cloud infrastructure, distributed systems, or large‑scale production environments.
  • Experience with Slurm and Linux administration.
  • Experience troubleshooting compute, storage, and networking systems.
Preferred Experience
  • GPU cluster operations.
  • NVIDIA technologies including CUDA, NCCL, NVLink, and GPUDirect.
  • InfiniBand networking.
  • Weka, Lustre, BeeGFS, or similar storage platforms.
  • Azure, AWS, or GCP.
  • Terraform, Ansible, or Infrastructure‑as‑Code.
  • PyTorch Distributed, Megatron‑LM, DeepSpeed, FSDP, or large‑scale AI training environments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Engineer – IFM
HPC Engineer – IFM

The Chronicle Of Higher Education, Inc. • United Arab Emirates

On-site
AED 120,000 - 180,000
Senior GPU HPC Engineer: Lead Frontier AI Infra
Senior GPU HPC Engineer: Lead Frontier AI Infra

The Chronicle Of Higher Education, Inc. • United Arab Emirates

On-site
High Performance Computing Software Engineer - Supercomputing
High Performance Computing Software Engineer - Supercomputing

Institute of Foundation Models • Abu Dhabi

On-site
AED 350,000 - 700,000
AI Infrastructure HPC Engineer for Large-Scale GPU Clusters
AI Infrastructure HPC Engineer for Large-Scale GPU Clusters

The Chronicle Of Higher Education, Inc. • United Arab Emirates

On-site
AED 120,000 - 180,000
Project Manager - High Performance Computing (HPC)
Project Manager - High Performance Computing (HPC)

Institute of Foundation Models • Abu Dhabi

On-site
Senior HPC Engineer: Build & Optimize AI Clusters
Senior HPC Engineer: Build & Optimize AI Clusters

Core42 • Abu Dhabi

On-site
AED 420,000 - 650,000
Competitive Salary
Yearly Bonus
Discount Cards Esaad and Fazaa
+2
Technical Program Manager, Compute Qualification
Technical Program Manager, Compute Qualification

Remotedxb • Dubai

On-site
AED 350,000 - 700,000
AI Infrastructure Engineer
AI Infrastructure Engineer

APPIT Software Inc. • Abu Dhabi

On-site
AED 180,000 - 240,000
Senior Fullstack Engineer
Senior Fullstack Engineer

Institute of Foundation Models • Abu Dhabi

On-site
AED 180,000 - 250,000
Senior Engineer - HPC Operations
Senior Engineer - HPC Operations

Core42 • Abu Dhabi

On-site
AED 350,000 - 650,000
Competitive Salary
Yearly Bonus
Exclusive Discount Cards: Esaad & Faza
+2