Senior HPC Engineer

Core42

Abu Dhabi

On-site

AED 420,000 - 650,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive Salary
Yearly Bonus
Discount Cards Esaad and Fazaa
Premium Family Insurance
Learning & Development

Job summary

Core42 is seeking a Senior HPC Engineer to design, deploy, and operate high-performance computing infrastructure across compute, storage, networking, and interconnects. You will lead deployment of HPC clusters, manage OpenSM, and optimize GPU-enabled workloads for AI/ML tasks.

You will work with Linux systems administration, job scheduling platforms (LSF preferred), and automation tools to ensure highly available HPC environments.

Qualifications

  • 5+ years in HPC operations, systems engineering, or related roles.
  • Strong Linux administration in enterprise data centers.
  • Hands-on experience deploying HPC clusters and managing interconnects.

Responsibilities

  • Design, implement, operate, and maintain Core42’s HPC infrastructure (compute, storage, networking).
  • Install, configure, deploy, and maintain HPC clusters and supporting infra.
  • Configure and troubleshoot InfiniBand and high-speed networks; manage OpenSM and routing.

Skills

HPC Operations
Linux Admin
InfiniBand
GPU workloads
Networking
Python scripting
Shell scripting
Troubleshooting
Monitoring
VMware ESXi

Education

Bachelor’s or Master’s in Computer Science, Engineering, Information Technology

Tools

LSF
OpenSM
xCAT
PXE/Kickstart
NVIDIA CUDA
Zabbix
Grafana

Job description

Core42, a leader in AI-powered cloud and digital infrastructure, is driving transformative technology solutions globally. Leveraging advanced resources and partnerships, Core42 empowers clients to harness sovereign AI infrastructure, especially in sectors with stringent regulatory needs. With a mission to redefine digital transformation, we combine sovereign capabilities with scalable, high-performance compute infrastructure, positioning itself at the forefront of AI innovation in the Middle East and beyond.

The opportunity

We are seeking a highly skilled Senior HPC Engineer to support the design, implementation, deployment, and ongoing operations of high-performance computing infrastructure. The role will be responsible for ensuring the availability, reliability, and performance of complex HPC environments spanning compute, storage, networking, InfiniBand, GPU resources, job scheduling, and supporting platforms.

The ideal candidate will bring strong hands‑on experience across Linux systems administration, HPC cluster deployment, network engineering, storage, workload management, and automation. You will work closely with internal engineering teams, customers, and technology vendors to troubleshoot complex issues, optimize infrastructure, and maintain highly available HPC environments supporting demanding workloads, including AI/ML and specialized industry applications.

Your key responsibilities
  • Support the design, implementation, operation, and maintenance of Core42’s HPC infrastructure, including compute, storage, networking, InfiniBand, and associated management platforms.
  • Install, configure, deploy, and maintain HPC clusters, including compute nodes, storage nodes, interconnects, and supporting infrastructure.
  • Configure, manage, and troubleshoot InfiniBand and high-speed Ethernet networks, including NVIDIA Mellanox switches, subnet managers such as OpenSM, and routing configurations.
  • Implement and maintain job scheduling and workload management platforms, with strong experience in LSF preferred and exposure to Slurm and/or PBS.
  • Administer and support Red Hat Enterprise Linux and Windows Server environments within complex HPC and data center infrastructures.
  • Support HPC storage environments, including IBM Spectrum Scale (GPFS), DDN GridScaler, NetApp SAN/NAS, and backup solutions such as IBM Spectrum Protect (TSM).
  • Utilize HPC administration and provisioning toolkits, including xCAT, PXE, Kickstart, and other automated deployment technologies.
  • Support and optimize GPU-enabled infrastructure, including NVIDIA H100 or newer architectures, and tune workloads for CPU and GPU resources.
  • Monitor system, network, storage, and infrastructure health using tools such as Zabbix, Grafana, and other monitoring and observability platforms.
  • Perform advanced troubleshooting and root cause analysis across complex HPC, network, storage, and operating system environments.
  • Support virtualization platforms including VMware ESXi/vSAN, Citrix, VxRail, and KVM where required.
  • Administer and support enterprise server platforms, including HPE ProLiant, Dell PowerEdge, and other compute infrastructure.
  • Develop automation and operational tooling using Shell scripting and Python to improve efficiency, reliability, and repeatability.
  • Create and maintain technical documentation, architecture diagrams, and operational procedures using tools such as Visio and Draw.io.
  • Apply HPC security best practices, including encryption, firewalls, access controls, network security, and compliance requirements.
  • Collaborate effectively with customers, vendors, and internal multidisciplinary teams to resolve technical issues and support infrastructure improvements.
  • Support specialized HPC workloads and, where applicable, industry applications such as reservoir engineering simulators and petroleum software tools including tNavigator, Nexus-VIP, Eclipse, Intersect, IMPOWER, and PUMA.
  • Participate in continuous improvement initiatives to enhance the performance, scalability, availability, and reliability of HPC environments.
Qualifications:
What we’re looking for
(a) Required skills / qualifications
  • Bachelor’s or Master’s degree in Computer Science, Engineering, Information Technology, or a related technical field.
  • 5+ years of experience in HPC operations, systems engineering, infrastructure engineering, network engineering, or a related technical role, preferably within an HPC environment.
  • Strong technical background with extensive hands‑on experience supporting multiple technologies and complex, highly available HPC environments.
  • Strong experience with Red Hat Enterprise Linux systems administration and enterprise data center infrastructure.
  • Hands‑on experience deploying and administering HPC clusters, including compute, storage, networking, and high‑speed interconnect components.
  • Strong expertise in InfiniBand networking and Ethernet technologies, including the configuration and management of NVIDIA Mellanox InfiniBand switches.
  • Experience with subnet managers such as OpenSM, routing configurations, network protocols, and troubleshooting methodologies within HPC environments.
  • Experience with HPC job scheduling and workload management systems; LSF experience is strongly preferred, with Slurm and/or PBS experience also highly desirable.
  • Familiarity with IBM Spectrum Scale (GPFS), IBM Spectrum Protect (TSM), DDN GridScaler, and NetApp SAN/NAS storage technologies.
  • Experience with HPC provisioning and administration tools such as xCAT, PXE, and Kickstart.
  • Knowledge of GPU infrastructure and workload optimization, with experience supporting NVIDIA H100 or newer GPU technologies considered desirable.
  • Experience tuning workloads and infrastructure for specific CPU and GPU hardware configurations.
  • Familiarity with VMware ESXi/vSAN, Citrix, VxRail, and KVM virtualization technologies.
  • Strong Windows Server administration experience, including exposure to Citrix VDI environments.
  • Experience supporting enterprise server hardware, including HPE ProLiant, Dell PowerEdge, and similar platforms.
  • Proficiency with monitoring and observability tools such as Zabbix, Grafana, or equivalent platforms.
  • Strong scripting and automation skills using Shell scripting and Python.
  • Strong understanding of LAN/WAN networking, including switches, routers, and associated network technologies.
  • Good knowledge of cybersecurity requirements, network security, firewalls, encryption, and HPC security best practices.
  • Familiarity with network monitoring tools and advanced troubleshooting techniques.
  • Strong communication, stakeholder management, and negotiation skills, with the ability to work effectively with customers, vendors, and internal teams.
  • Experience creating and maintaining technical diagrams and documentation using Visio, Draw.io, or similar tools.
  • Familiarity with reservoir engineering simulators and petroleum software, including tNavigator, Nexus-VIP, Schlumberger Eclipse & Intersect, ExxonMobil IMPOWER, Total PUMA, or similar platforms, is an advantage.
What working at Core42 offers

With a diverse team of 1,100+ employees from 68 nationalities, we foster an inclusive, innovative and collaborative environment. At Core42, we foster a culture grounded in trust, accountability and high performance. We are united by our values: Grit, where we overcome challenges with resilience and determination, Passion, which drives us to pursue excellence in everything we do, and Impact, as we aim to inspire progress and create meaningful change. Our team members thrive in an environment where each person’s contributions propel us forward, and together, we commit to achieving extraordinary results.

  • Competitive Salary: We offer an attractive salary package based on your skills and experience.
  • Yearly Bonus: In recognition of your contributions, you will receive a performance-based annual bonus.
  • Exclusive Discount Cards: Access special benefits with Esaad and Fazaa cards, offering discounts across a wide range of services.
  • Premium Family Insurance: We provide comprehensive health coverage, including dental, vision, and life insurance, ensuring the well‑being of you and your family.
  • Learning & Development: We offer access to top‑tier learning platforms to help you grow in your career. Learn at your own pace with unlimited access to premium courses.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Engineer - HPC Operations
Senior Engineer - HPC Operations

Core42 • Abu Dhabi

On-site
AED 350,000 - 650,000
Competitive Salary
Yearly Bonus
Exclusive Discount Cards: Esaad & Faza
+2
Senior Engineer - Infrastructure and Cloud Engineering
Senior Engineer - Infrastructure and Cloud Engineering

Core42 • Abu Dhabi

On-site
AED 320,000 - 520,000
Yearly Bonus
Exclusive Discount Cards: Esaad and FZ
Premium Family Insurance
Senior Software Engineer
Senior Software Engineer

Core42 • Dubai

On-site
AED 350,000 - 500,000
Yearly bonus
Discount cards (Esaad/Fazaa)
Premium family insurance
+1
Senior Architect - Storage and Data Protection
Senior Architect - Storage and Data Protection

Core42 • Abu Dhabi

On-site
AED 450,000 - 750,000
Competitive Salary
Yearly Bonus
Exclusive Discount Cards
+2
Senior DevOps Engineer
Senior DevOps Engineer

Core42 • Abu Dhabi

On-site
AED 300,000 - 520,000
Competitive Salary
Yearly Bonus
Exclusive Discount Cards
+2
Senior Software Engineer (Go)
Senior Software Engineer (Go)

Core42 • Abu Dhabi

On-site
AED 320,000 - 520,000
Competitive Salary
Yearly Bonus
Discount Cards (Esaad & Fazaa)
+2
Senior DevOps Engineer
Senior DevOps Engineer

Core42 • United Arab Emirates

On-site
AED 350,000 - 600,000
Competitive Salary
Yearly Bonus
Discount Cards
+2
Senior Manager - Private Cloud Sales
Senior Manager - Private Cloud Sales

Core42 • Dubai

On-site
AED 661,000 - 992,000
Yearly Bonus
Esaad/Fazaa discounts
Premium health insurance
+1
Senior Engineer - Container Platforms
Senior Engineer - Container Platforms

Core42 • United Arab Emirates

On-site
AED 300,000 - 420,000
Competitive salary
Yearly bonus
Discount cards (Esaad/Fazaa)
+2
Specialist - Information Security
Specialist - Information Security

Core42 • Abu Dhabi

On-site
AED 520,000 - 760,000
Competitive salary
Yearly bonus
Discount cards