Senior HPC & GPU Cluster Engineer — Remote EU

Verda

España

On-site

PHP 6,536,000 - 9,441,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity included
Healthcare
Lunch
Wellbeing programs

Job summary

Verda in Helsinki is building a global AI cloud and is seeking an experienced HPC Engineer to own baremetal and virtualized GPU clusters behind our AI cloud. You will manage the InfiniBand fabric, shared filesystems, and workload orchestration to keep clusters healthy and ready for the next workload.

The role requires deep Linux knowledge, InfiniBand expertise, Slurm experience, and scripting, with a path to impact across a fast-growing, remote-friendly organization offering equity and

Qualifications

  • Experience administering baremetal GPU/HPC clusters end to end.
  • Experience with virtualization and GPU virtualization stacks.
  • Familiarity with shared parallel filesystems (Lustre/GPFS).

Responsibilities

  • Administer baremetal GPU/HPC clusters end to end, from provisioning through day-two operations
  • Administer virtualized clusters, including the hypervisor and GPU virtualization stack
  • Design, deploy, and tune InfiniBand fabrics, including topology planning, subnet management, and performance validation
  • Deploy and operate shared/parallel filesystems supporting training and inference workloads, balancing performance, capacity, and reliability
  • Troubleshoot and resolve issues across the full fabric stack: fibers, transceivers, NICs, switches, drivers, and firmware
  • Partner with remote-hands and data center teams to diagnose hardware faults and execute physical-layer fixes and cluster expansions
  • Operate and tune Slurm (or equivalent) workload scheduling deployments used by customers and internal teams
  • Keep issue tracking, IPAM, and DCIM records accurate as clusters are built, changed, and decommissioned
  • Participate in on-call rotations and incident response for cluster-level issues
  • Collaborate with platform, network, and storage teams to integrate new clusters into the broader AI cloud

Skills

Linux expertise
Memory management
InfiniBand networking
NVIDIA CUDA/NCCL stack
Slurm workload scheduling
Scripting (Python, Bash)
Cluster health & performance

Tools

SLURM
Lustre/GPFS
InfiniBand fabric tooling
KVM/QEMU GPU passthrough

Job description

Verda in Helsinki is building a global AI cloud and is seeking an experienced HPC Engineer to own baremetal and virtualized GPU clusters behind our AI cloud. You will manage the InfiniBand fabric, shared filesystems, and workload orchestration to keep clusters healthy and ready for the next workload.

The role requires deep Linux knowledge, InfiniBand expertise, Slurm experience, and scripting, with a path to impact across a fast-growing, remote-friendly organization offering equity and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior GPU HPC Systems Architect for AI Cloud
Senior GPU HPC Systems Architect for AI Cloud

Lambda • Santo Niño 1st

Hybrid
PHP 11,942,000 - 15,085,000
Health, dental, vision coverage foryou
Wellness and commuter stipends
401k Plan with 2% company match (USA)
+1
Data Center Engineer/Lead
Data Center Engineer/Lead

Atomic Recruitment SEA • Metro Manila

On-site
PHP 391,000 - 670,000
Remote Cloud Performance & Cost Engineer
Remote Cloud Performance & Cost Engineer

BairesDev • Mexico

On-site
PHP 5,474,000 - 9,124,000
Remote work
USD compensation
Home setup
+4
Lead HPC Network Architect for AI Cloud Infra
Lead HPC Network Architect for AI Cloud Infra

Lambda • Santo Niño 1st

On-site
PHP 11,314,000 - 16,342,000
Health, dental, and vision coverage
401k with 2% company match
Wellness and commuter stipends
+1
Senior DevOps Engineer — Remote Europe, Cloud Infra & CI/CD
Senior DevOps Engineer — Remote Europe, Cloud Infra & CI/CD

Bet On Talent • Hinoba-an

Hybrid
PHP 7,524,000 - 10,031,000
28 days off + public holidays
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

NVIDIA Corporation • Hinoba-an

On-site
PHP 2,000,000 - 4,500,000
Hybrid work model
Staff HPC Network Architect
Staff HPC Network Architect

Lambda • Santo Niño 1st

On-site
PHP 11,314,000 - 16,342,000
Health, dental, and vision coverage
401k with 2% company match
Wellness and commuter stipends
+1
Data Centre Field Operations Engineer
Data Centre Field Operations Engineer

Nava • Metro Manila

On-site
PHP 1,200,000 - 2,000,000
GPU Data Center Ops Lead | On-Site Incident Command
GPU Data Center Ops Lead | On-Site Incident Command

Nava • Metro Manila

On-site
PHP 1,200,000 - 2,000,000
Staff HPC Systems Architect
Staff HPC Systems Architect

Lambda • Santo Niño 1st

On-site
PHP 11,942,000 - 15,085,000
Health, dental, vision coverage foryou
Wellness and commuter stipends
401k Plan with 2% company match (USA)
+1