System Engineer

Hermes Corporate

Brussel

Sur place

EUR 85 000 - 120 000

Plein temps

Il y a 2 jours
Soyez parmi les premiers à postuler
Générateur de candidature

Transformez ce poste en entretien — un CV et une lettre de motivation conçus selon ce que cet employeur recherche.

Passez les filtres ATS

Résumé du poste

Hermes Corporate is seeking an experienced AI & HPC Infrastructure Engineer to operate, optimise, and evolve high-performance computing infrastructure supporting advanced engineering and scientific workloads.

This hands-on role requires strong Linux and HPC expertise, ownership of production platforms, and the ability to troubleshoot complex systems across GPU compute, storage, networking, containers, and platform performance.

Qualifications

  • Substantial professional experience with Linux systems engineering at an advanced level.
  • Proven track record operating production HPC environments and large-scale workloads.
  • Hands-on experience administering NVIDIA-based GPU compute infrastructure.
  • Experience with containerisation in HPC, including multi-node workloads.
  • Strong troubleshooting, performance tuning, and ownership of production infrastructure.
  • Professional working proficiency in English.

Responsabilités

  • Administer and maintain GPU compute infrastructure, including system configuration and drivers.
  • Operate, monitor, and optimise HPC clusters for reliability and performance.
  • Define resource allocation policies across workload types for fair use.
  • Manage containerised environments, provisioning, and GPU access.
  • Administer shared storage and support high-speed interconnects.
  • Lead performance engineering: profiling, benchmarking, bottleneck analysis.
  • Support integration and optimisation of applications on GPU/HPC platforms.
  • Provide incident response, troubleshooting, and root-cause analysis.
  • Maintain platform monitoring, alerts, and operational reporting.
  • Develop technical documentation, procedures, and platform standards.
  • Contribute to security, architecture governance, and compliance activities.

Connaissances

Linux systems engineering
HPC cluster administration
NVIDIA GPU compute
Containerisation
Performance tuning
Incident response
English proficiency

Outils

NVIDIA drivers
GPU software stack
Container runtimes
Monitoring tools

Description du poste

We are looking for an experienced AI & HPC Infrastructure Engineer to operate, optimise, and continuously evolve high-performance computing infrastructure supporting advanced engineering and scientific workloads.

This is a hands-on infrastructure role for someone with strong Linux and HPC expertise who is comfortable owning production platforms, troubleshooting complex systems, and driving improvements across GPU compute, storage, networking, containers, and platform performance.

What You’ll Do
  • Administer and maintain GPU compute infrastructure, including system configuration, NVIDIA drivers and software stack, firmware, and hardware health monitoring.
  • Operate, monitor, and optimise HPC clusters to ensure reliability, performance, and efficient use of compute resources.
  • Define and refine resource allocation policies across different workload types to ensure effective and equitable use of available capacity.
  • Manage containerised environments for platform users, including user provisioning, environment maintenance, GPU access, and standards for container usage.
  • Administer shared and high-performance storage and support the operation and troubleshooting of high-speed interconnects.
  • Lead performance engineering activities, including profiling, benchmarking, bottleneck analysis, and platform optimisation.
  • Support the integration and optimisation of engineering and scientific applications on GPU-accelerated and HPC platforms.
  • Provide incident response, troubleshooting, root-cause analysis, and change management, including planning and execution of maintenance activities.
  • Maintain platform monitoring, alerting, and operational reporting.
  • Develop and maintain technical documentation, operational procedures, and platform standards.
  • Contribute to security, architecture governance, and compliance activities.
Required Qualifications
  • Substantial professional experience in Linux systems engineering / administration at an advanced level.
  • Proven experience operating production HPC environments, including cluster administration and large-scale parallel workloads.
  • Hands-on experience administering GPU compute infrastructure, particularly NVIDIA-based environments and the associated software stack.
  • Practical experience with containerisation in HPC environments, including GPU device access and multi-node workloads.
  • Strong troubleshooting and performance-tuning skills across complex compute infrastructure.
  • Demonstrated ability to take end-to-end ownership of production infrastructure and act as a senior technical escalation point.
  • Strong analytical and problem-solving skills with a production-focused mindset.
  • Professional working proficiency in English.
Preferred Qualifications
  • Experience supporting engineering simulation, scientific computing, or computational analysis workloadson accelerated infrastructure.
  • Experience with HPC scheduling and resource management.
  • Experience with monitoring and observability for compute infrastructure.
  • Experience with high-performance storage and networking/interconnect technologies.
  • Experience in an enterprise, industrial, or regulated environment.
  • Experience with platform automation, scripting, or infrastructure tooling.
Ideal Candidate

The ideal candidate is a hands-on infrastructure engineer who combines deep Linux, HPC, and GPU expertise with strong operational ownership.

You are comfortable working close to the hardware and software stack, diagnosing complex performance and reliability issues, and improving infrastructure used by demanding engineering and scientific workloads. You thrive in environments where reliability, performance, and technical depth matter, and you can operate effectively as a senior technical point of reference for the platform.

Obtenez votre examen gratuit et confidentiel de votre CV.

ou faites glisser et déposez votre fichier ici.

Similar jobs

Postes similaires à comparer

Senior AI & HPC Infrastructure Engineer
Senior AI & HPC Infrastructure Engineer

Hermes Corporate • Brussel

Sur place
EUR 85 000 - 120 000
HPC/Kubernetes SW Engineer
HPC/Kubernetes SW Engineer

Keysight Technologies SAles Spain SL. • Oost-Vlaanderen

Sur place
EUR 70 000 - 110 000
HPC/Kubernetes Software Engineer
HPC/Kubernetes Software Engineer

Keysight Technologies • Oost-Vlaanderen

Sur place
EUR 65 000 - 90 000
Senior HPC GPU Cluster Engineer – Flexible, Global Impact
Senior HPC GPU Cluster Engineer – Flexible, Global Impact

Nebius Group • Brussel

Sur place
EUR 90 000 - 140 000
Competitive pay
Career growth
Flexibility
+3
Tech Ops Engineer
Tech Ops Engineer

Haulogy • Waals-Brabant

Hybride
EUR 42 000 - 68 000
Senior HPC Cluster Engineer
Senior HPC Cluster Engineer

Nebius Group • Brussel

Sur place
EUR 90 000 - 140 000
Competitive pay
Career growth
Flexibility
+3
Senior Data Scientist & HPC Cloud Engineer
Senior Data Scientist & HPC Cloud Engineer

Altia • Brussel

Sur place
EUR 90 000 - 130 000
High-impact project
Collaborative work environment
Career Plan
+2
Cloud Platform Engineer
Cloud Platform Engineer

act digital • Brussel Hoofdstad

Sur place
EUR 70 000 - 110 000
Cloud Engineer
Cloud Engineer

Fusit • Oost-Vlaanderen

Sur place
EUR 85 000 - 120 000
Cloud Engineer
Cloud Engineer

Fusit • Gent

Sur place
EUR 85 000 - 120 000