Senior HPC Operations Engineer for AI/ML Workloads

Core42

Abu Dhabi

On-site

AED 350,000 - 650,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive Salary
Yearly Bonus
Exclusive Discount Cards: Esaad & Faza
Premium Family Insurance
Learning & Development access

Job summary

Core42, a leader in AI-powered cloud and digital infrastructure, seeks a Senior Engineer – HPC Operations to oversee daily operations of large-scale AI/ML HPC clusters. You will ensure stable, secure, high-performance infrastructure using Slurm, Kubernetes, and modern MLOps tools.

The role requires deep HPC experience, scripting proficiency, and strong collaboration across global teams, with mentorship to engineers and on-call participation as needed.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
  • 7+ years of experience in HPC operations, systems engineering, or DevOps roles.
  • Hands-on experience configuring and maintaining complex HPC environments.
  • Experience with Slurm clusters and Kubernetes-based AI/ML workloads.
  • GPU resource management and performance tuning for AI/ML workloads.
  • Monitoring and observability using Prometheus, Grafana, DCGM.
  • Strong scripting and automation skills (Python, Bash, Ansible, Terraform).
  • In-depth Linux, networking, and storage knowledge (NFS, Lustre, Ceph).

Responsibilities

  • Lead daily operational support of HPC infrastructure including compute, storage, networking, and schedulers (Slurm, Kubernetes).
  • Maximize efficiency and performance of HPC systems with optimal resource utilization and minimal downtime.
  • Serve as primary technical escalation point for L2 support and incidents.
  • Monitor health and performance using tools like Prometheus, Grafana, DCGM.
  • Manage user environments for AI/ML workloads with containers and workflow tools.
  • Implement and manage job scheduling policies and partitions for fairness.
  • Lead root cause analysis and document post-mortems and improvements.
  • Mentor junior engineers and participate in on-call rotation.
  • Ensure security and policy compliance; assist audits and changes.

Skills

HPC operations
Scripting & automation
Linux systems expertise
GPU resource management
AI/ML workload optimization

Education

Bachelor’s or Master’s degree in CS/Engineering or related field

Tools

Slurm
Kubernetes
Prometheus
Grafana
DCGM
NFS
Lustre
Ceph
RDMA networking
InfiniBand and RoCE

Job description

Core42, a leader in AI-powered cloud and digital infrastructure, seeks a Senior Engineer – HPC Operations to oversee daily operations of large-scale AI/ML HPC clusters. You will ensure stable, secure, high-performance infrastructure using Slurm, Kubernetes, and modern MLOps tools.

The role requires deep HPC experience, scripting proficiency, and strong collaboration across global teams, with mentorship to engineers and on-call participation as needed.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior HPC Engineer: Build & Optimize AI Clusters
Senior HPC Engineer: Build & Optimize AI Clusters

Core42 • Abu Dhabi

On-site
AED 420,000 - 650,000
Competitive Salary
Yearly Bonus
Discount Cards Esaad and Fazaa
+2
Senior Engineer - HPC Operations
Senior Engineer - HPC Operations

Core42 • Abu Dhabi

On-site
AED 350,000 - 650,000
Competitive Salary
Yearly Bonus
Exclusive Discount Cards: Esaad & Faza
+2
Senior HPC Engineer
Senior HPC Engineer

Core42 • Abu Dhabi

On-site
AED 420,000 - 650,000
Competitive Salary
Yearly Bonus
Discount Cards Esaad and Fazaa
+2
Senior Container Platform Engineer — Kubernetes & OpenShift
Senior Container Platform Engineer — Kubernetes & OpenShift

Core42 • United Arab Emirates

On-site
AED 300,000 - 420,000
Competitive salary
Yearly bonus
Discount cards (Esaad/Fazaa)
+2
Senior Data Platform Engineer: Kubernetes & Cloud
Senior Data Platform Engineer: Kubernetes & Cloud

Core42 • Dubai

On-site
AED 480,000 - 720,000
Yearly bonus
Exclusive discount cards
Premium family insurance
Senior HPC Engineer – IFM
Senior HPC Engineer – IFM

The Chronicle Of Higher Education, Inc. • United Arab Emirates

On-site
Senior Linux & Kubernetes Engineer for HPC Infra
Senior Linux & Kubernetes Engineer for HPC Infra

DeepSource Technologies • United Arab Emirates

On-site
AED 180,000 - 280,000
Senior Java Backend Engineer - Cloud & OpenStack Expert
Senior Java Backend Engineer - Cloud & OpenStack Expert

Core42 • Dubai

On-site
AED 350,000 - 500,000
Yearly bonus
Discount cards (Esaad/Fazaa)
Premium family insurance
+1
Senior DevOps Engineer — Cloud, CI/CD & Kubernetes Leader
Senior DevOps Engineer — Cloud, CI/CD & Kubernetes Leader

Core42 • United Arab Emirates

On-site
AED 350,000 - 600,000
Competitive Salary
Yearly Bonus
Discount Cards
+2
Senior Engineer - Data Platforms
Senior Engineer - Data Platforms

Core42 • Dubai

On-site
AED 480,000 - 720,000
Yearly bonus
Exclusive discount cards
Premium family insurance