Senior Engineer - HPC Operations

Core42

Abu Dhabi

On-site

AED 350,000 - 650,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive Salary
Yearly Bonus
Exclusive Discount Cards: Esaad & Faza
Premium Family Insurance
Learning & Development access

Job summary

Core42, a leader in AI-powered cloud and digital infrastructure, seeks a Senior Engineer – HPC Operations to oversee daily operations of large-scale AI/ML HPC clusters. You will ensure stable, secure, high-performance infrastructure using Slurm, Kubernetes, and modern MLOps tools.

The role requires deep HPC experience, scripting proficiency, and strong collaboration across global teams, with mentorship to engineers and on-call participation as needed.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
  • 7+ years of experience in HPC operations, systems engineering, or DevOps roles.
  • Hands-on experience configuring and maintaining complex HPC environments.
  • Experience with Slurm clusters and Kubernetes-based AI/ML workloads.
  • GPU resource management and performance tuning for AI/ML workloads.
  • Monitoring and observability using Prometheus, Grafana, DCGM.
  • Strong scripting and automation skills (Python, Bash, Ansible, Terraform).
  • In-depth Linux, networking, and storage knowledge (NFS, Lustre, Ceph).

Responsibilities

  • Lead daily operational support of HPC infrastructure including compute, storage, networking, and schedulers (Slurm, Kubernetes).
  • Maximize efficiency and performance of HPC systems with optimal resource utilization and minimal downtime.
  • Serve as primary technical escalation point for L2 support and incidents.
  • Monitor health and performance using tools like Prometheus, Grafana, DCGM.
  • Manage user environments for AI/ML workloads with containers and workflow tools.
  • Implement and manage job scheduling policies and partitions for fairness.
  • Lead root cause analysis and document post-mortems and improvements.
  • Mentor junior engineers and participate in on-call rotation.
  • Ensure security and policy compliance; assist audits and changes.

Skills

HPC operations
Scripting & automation
Linux systems expertise
GPU resource management
AI/ML workload optimization

Education

Bachelor’s or Master’s degree in CS/Engineering or related field

Tools

Slurm
Kubernetes
Prometheus
Grafana
DCGM
NFS
Lustre
Ceph
RDMA networking
InfiniBand and RoCE

Job description

Core42, a leader in AI-powered cloud and digital infrastructure, is driving transformative technology solutions globally. Leveraging advanced resources and partnerships, Core42 empowers clients to harness sovereign AI infrastructure, especially in sectors with stringent regulatory needs. With a mission to redefine digital transformation, we combine sovereign capabilities with scalable, high-performance compute infrastructure, positioning itself at the forefront of AI innovation in the Middle East and beyond.

The opportunity

We are seeking a highly skilled Senior Engineer – HPC Operations to oversee the daily operations and support of high-performance computing clusters designed to power large-scale AI and ML workloads. This role ensures stable, secure, and high-performing infrastructure leveraging technologies such as Slurm, Kubernetes, and modern MLOps platforms. The ideal candidate will bring deep technical expertise in HPC and a strong operational mindset to drive continuous improvement and automation across globally distributed environments. Responsibilities will extend to collaborating with multidisciplinary teams, leading complex projects, implementing cutting-edge technologies, and providing mentorship to operations engineers.

Your key responsibilities
  • Lead the daily operational support of HPC infrastructure including compute, storage, networking, and scheduler components (Slurm, Kubernetes, etc.).
  • Lead efforts to maximize the efficiency and performance of HPC systems, ensuring optimal resource utilization and minimal downtime.
  • Act as the primary technical escalation point for L2 support teams and ensure prompt resolution of incidents and service requests.
  • Monitor system health, performance, and utilization using advanced tools (e.g., Prometheus, Grafana, DCGM).
  • Manage user environments for AI/ML workloads including container orchestration (e.g., Docker, Kubernetes) and workflow tools (e.g., MLflow, Kubeflow).
  • Implement and manage job scheduling policies, priorities, and partitions within Slurm and/or Kubernetes environments to ensure fairness and efficiency.
  • Lead root cause analysis (RCA) of operational issues and contribute to post-mortem documentation and continuous improvement efforts.
  • Provide mentorship and guidance to junior engineers and participate in on-call rotation if required.
  • Ensure compliance with security and operational policies; assist in audits and documentation for change and incident management processes.
Qualifications:
What we’re looking for
(a) Required skills / qualifications
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related technical field.
  • 7+ years of experience in HPC operations, systems engineering, or DevOps roles.
  • Advanced knowledge and expertise in configuring, optimizing, and maintaining complex HPC environments, including hardware, software, and storage systems.
  • Hands-on experience managing Slurm clusters and/or Kubernetes-based environments for AI/ML workloads.
  • Expert knowledge of GPU resource management, workload schedulers, and performance tuning for AI/ML workloads.
  • Experience with monitoring and observability frameworks such as Prometheus, Grafana, and DCGM.
  • Strong scripting and automation skills (Python, Bash, Ansible, Terraform).
  • In-depth understanding of Linux (RHEL/CentOS/Ubuntu), networking concepts (RDMA, InfiniBand, RoCE), and storage technologies (NFS, Lustre, Ceph).
What working at Core42 offers

With a diverse team of 1,100+ employees from 68 nationalities, we foster an inclusive, innovative and collaborative environment. At Core42, we foster a culture grounded in trust, accountability and high performance. We are united by our values: Grit, where we overcome challenges with resilience and determination, Passion, which drives us to pursue excellence in everything we do, and Impact, as we aim to inspire progress and create meaningful change. Our team members thrive in an environment where each person’s contributions propel us forward, and together, we commit to achieving extraordinary results.

  • Competitive Salary: We offer an attractive salary package based on your skills and experience
  • Yearly Bonus: In recognition of your contributions, you will receive a performance-based annual bonus
  • Exclusive Discount Cards: Access special benefits with Esaad and Fazaa cards, offering discounts across a wide range of services
  • Premium Family Insurance: We provide comprehensive health coverage, including dental, vision and life insurance, ensuring the well-being of you and your family
  • Learning & Development: We offer access to top-tier learning platforms to help you grow in your career. Learn at your own pace with unlimited access to premium courses.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer
Senior DevOps Engineer

Core42 • United Arab Emirates

On-site
AED 350,000 - 600,000
Competitive Salary
Yearly Bonus
Discount Cards
+2
Senior Software Engineer
Senior Software Engineer

Core42 • Dubai

On-site
AED 350,000 - 500,000
Yearly bonus
Discount cards (Esaad/Fazaa)
Premium family insurance
+1
Senior Devops Engineer
Senior Devops Engineer

Core42 • Dubai

On-site
AED 320,000 - 520,000
Yearly Bonus
Exclusive Discount Cards: Esaad and Fa
Premium Family Insurance
+1
Senior Software Engineer (Go)
Senior Software Engineer (Go)

Core42 • Abu Dhabi

On-site
AED 320,000 - 520,000
Competitive Salary
Yearly Bonus
Discount Cards (Esaad & Fazaa)
+2
Senior Engineer - Data Platforms
Senior Engineer - Data Platforms

Core42 • Dubai

On-site
AED 480,000 - 720,000
Yearly bonus
Exclusive discount cards
Premium family insurance
Senior HPC Operations Engineer for AI/ML Workloads
Senior HPC Operations Engineer for AI/ML Workloads

Core42 • Abu Dhabi

On-site
AED 350,000 - 650,000
Competitive Salary
Yearly Bonus
Exclusive Discount Cards: Esaad & Faza
+2
Senior Manager - Private Cloud Sales
Senior Manager - Private Cloud Sales

Core42 • Dubai

On-site
AED 661,000 - 992,000
Yearly Bonus
Esaad/Fazaa discounts
Premium health insurance
+1
Specialist - Information Security
Specialist - Information Security

Core42 • Abu Dhabi

On-site
AED 520,000 - 760,000
Competitive salary
Yearly bonus
Discount cards
Security Engineer
Security Engineer

Core42 • United Arab Emirates

On-site
AED 350,000 - 550,000
Yearly bonus
Competitive salary
Exclusive discount cards (Esaad Faza)
+2
Specialist - Information Security
Specialist - Information Security

Core42 • Dubai

On-site
AED 380,000 - 600,000
Competitive Salary
Yearly Bonus
Discount Cards (Esaad/Fazaa)
+2