HPC ENGINEER

Hermes Corporate

Italia

In loco

EUR 70.000 - 120.000

Tempo pieno

2 giorni fa
Candidati tra i primi

Ricevi più risposte dai datori di lavoro

Invia un CV specifico per questa offerta in pochi minuti.

Descrizione del lavoro

Hermes Corporate seeks an experienced AI & HPC Infrastructure Engineer to operate, optimise and evolve high-performance computing platforms that power advanced engineering and scientific workloads.

This hands-on role requires strong Linux and HPC expertise, ownership of production platforms, and the ability to troubleshoot complex systems across GPU compute, storage, networking, containers, and platform performance.

Competenze

  • Proven experience in Linux systems engineering at an advanced level.
  • Proficient in operating production HPC environments and large-scale parallel workloads.
  • Hands-on administration of GPU compute infrastructure and NVIDIA software stacks.
  • Experience with containerisation in HPC environments and multi-node workloads.
  • Strong debugging, profiling, and performance-tuning skills in complex compute infra.
  • Ability to own production infrastructure and serve as senior escalation point.

Mansioni

  • Administer GPU compute infrastructure including drivers, software stack, firmware, and hardware health.
  • Operate and optimise HPC clusters for reliability and performance.
  • Refine resource allocation policies across workloads for efficient capacity use.
  • Manage containerised environments for platform users, including provisioning and GPU access.
  • Maintain shared high-speed storage and support interconnects operation and troubleshooting.
  • Lead performance engineering activities: profiling, benchmarking, bottleneck analysis, platform optimisation.
  • Support integration and optimisation of engineering and scientific applications on GPU-accelerated platforms.
  • Provide incident response, root-cause analysis, and change management for maintenance planning.
  • Maintain platform monitoring, alerting, and operational reporting.
  • Develop and maintain technical docs, procedures, and platform standards.
  • Contribute to security, architecture governance, and compliance activities.

Conoscenze

Linux systems engineering
HPC cluster administration
NVIDIA GPU compute stack
Containerisation in HPC
Performance troubleshooting
End-to-end ownership

Strumenti

NVIDIA drivers
CUDA toolkit
Container tooling (Docker/Singularity)
Storage & networking tech

Descrizione del lavoro

We are looking for an experienced AI & HPC Infrastructure Engineer to operate, optimise, and continuously evolve high-performance computing infrastructure supporting advanced engineering and scientific workloads.

This is a hands-on infrastructure role for someone with strong Linux and HPC expertise who is comfortable owning production platforms, troubleshooting complex systems, and driving improvements across GPU compute, storage, networking, containers, and platform performance.

What You’ll Do
  • Administer and maintain GPU compute infrastructure, including system configuration, NVIDIA drivers and software stack, firmware, and hardware health monitoring.
  • Operate, monitor, and optimise HPC clusters to ensure reliability, performance, and efficient use of compute resources.
  • Define and refine resource allocation policies across different workload types to ensure effective and equitable use of available capacity.
  • Manage containerised environments for platform users, including user provisioning, environment maintenance, GPU access, and standards for container usage.
  • Administer shared and high-performance storage and support the operation and troubleshooting of high-speed interconnects.
  • Lead performance engineering activities, including profiling, benchmarking, bottleneck analysis, and platform optimisation.
  • Support the integration and optimisation of engineering and scientific applications on GPU-accelerated and HPC platforms.
  • Provide incident response, troubleshooting, root-cause analysis, and change management, including planning and execution of maintenance activities.
  • Maintain platform monitoring, alerting, and operational reporting.
  • Develop and maintain technical documentation, operational procedures, and platform standards.
  • Contribute to security, architecture governance, and compliance activities.
Required Qualifications
  • Substantial professional experience in Linux systems engineering / administration at an advanced level.
  • Proven experience operating production HPC environments, including cluster administration and large-scale parallel workloads.
  • Hands-on experience administering GPU compute infrastructure, particularly NVIDIA-based environments and the associated software stack.
  • Practical experience with containerisation in HPC environments, including GPU device access and multi-node workloads.
  • Strong troubleshooting and performance-tuning skills across complex compute infrastructure.
  • Demonstrated ability to take end-to-end ownership of production infrastructure and act as a senior technical escalation point.
  • Strong analytical and problem-solving skills with a production-focused mindset.
  • Professional working proficiency in English.
Preferred Qualifications
  • Experience supporting engineering simulation, scientific computing, or computational analysis workloadson accelerated infrastructure.
  • Experience with HPC scheduling and resource management.
  • Experience with monitoring and observability for compute infrastructure.
  • Experience with high-performance storage and networking/interconnect technologies.
  • Experience in an enterprise, industrial, or regulated environment.
  • Experience with platform automation, scripting, or infrastructure tooling.
Ideal Candidate

The ideal candidate is a hands-on infrastructure engineer who combines deep Linux, HPC, and GPU expertise with strong operational ownership.

You are comfortable working close to the hardware and software stack, diagnosing complex performance and reliability issues, and improving infrastructure used by demanding engineering and scientific workloads. You thrive in environments where reliability, performance, and technical depth matter, and you can operate effectively as a senior technical point of reference for the platform.

Ottieni la revisione del curriculum gratis e riservata.
o trascina qui il file.
Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Senior GPU HPC & AI Infrastructure Engineer
Senior GPU HPC & AI Infrastructure Engineer

Hermes Corporate • Italia

In loco
EUR 70.000 - 120.000
Infrastructure Engineer
Infrastructure Engineer

European Tech Recruit • Millan

In loco
EUR 50.000 - 70.000
Cloud / DevOps Engineer (AI Infrastructure)
Cloud / DevOps Engineer (AI Infrastructure)

AI4I Foundation • Piemonte

Ibrido
EUR 60.000 - 80.000
Competitive compensation
Collaborative environment
Access to advanced computing infrastructure
Infiniband Network Engineer
Infiniband Network Engineer

NVIDIA ITALY S.R.L. • Italia

In loco
EUR 50.000 - 114.000
Senior Linux System Administrator
Senior Linux System Administrator

Do IT Now • Emilia-Romagna

Ibrido
EUR 60.000 - 90.000
Birthday off
Professional development
Remote work possibility
+1
Senior AI Compute Engineer
Senior AI Compute Engineer

NVIDIA Gruppe • Italia

In loco
EUR 66.000 - 114.000
HPC / AI Specialist
HPC / AI Specialist

AI4I Foundation • Piemonte

Ibrido
EUR 70.000 - 90.000
Competitive compensation
Access to advanced computing infrastructure
Collaborative environment with leading researchers
Senior Service Delivery Manager / HPC–AI Application Specialist
Senior Service Delivery Manager / HPC–AI Application Specialist

AI4I • Torino

In loco
EUR 70.000 - 90.000
Competitive compensation
Flexible work arrangements
Leadership over critical AI/HPC services
+1
Linux Systems Administrator - HPC Simulations
Linux Systems Administrator - HPC Simulations

Allegro MicroSystems, LLC • Milano

Ibrido
EUR 55.000 - 85.000
Hybrid work model
Senior Service Delivery Manager / HPC–AI Application Specialist
Senior Service Delivery Manager / HPC–AI Application Specialist

AI4I Foundation • Piemonte

Ibrido
EUR 70.000 - 90.000
Competitive compensation
Flexible work arrangements
Access to advanced AI computing infrastructure