Senior HPC Network Engineer – RDMA & GPU Clusters

Multicoin

London

On-site

NOK 1,141,000 - 1,648,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary and equity sign-on
Biannual bonus scheme
Fully expensed tech to match your need
Breakfast and dinner allowance for off

Job summary

Fuse Energy is building a best-in-class network fabric for a multi-tenant AI cluster, spanning high-performance compute, storage, and management networks. You will own the fabric architecture through day-2 operations, including RDMA traffic and tenant isolation.

You’ll design, deploy, and operate leaf-spine data centre fabrics, automate provisioning, and build observability dashboards to detect congestion before tenants notice.

Qualifications

  • 5+ years as a network engineer operating production data centre networks.
  • Strong dynamic routing experience, esp. BGP; EVPN/VXLAN.
  • Hands-on leaf-spine / Clos fabric design and operation.
  • Proficient in Linux networking stack and modern data centre OS.
  • RDMA fabric experience: lossless Ethernet or InfiniBand.
  • Network automation with Python/Ansible and config-as-code.
  • Solid Linux administration fundamentals and telemetry proficiency.
  • Experience with Prometheus/Grafana and related telemetry.
  • Experience with campus networks, NAC/802.1X, VPN access.
  • Clear communicator who can teach and document.

Responsibilities

  • Design and operate lossless, RDMA-capable fabrics (e.g. RoCEv2, InfiniBand) for GPU compute and storage traffic, including QoS, congestion control, and buffer tuning at scale
  • Build and manage leaf-spine data centre fabrics, with routed underlay and overlay design (e.g. BGP, EVPN/VXLAN)
  • Implement and maintain per-tenant network isolation across compute, storage, and management planes.
  • Automate network provisioning, configuration, and validation, treating switch config as code (e.g. Ansible, Python, NetBox as source of truth), deployed through CI
  • Build telemetry and observability for the fabric: flow-level and buffer-level visibility, dashboards, and alerting that catches congestion and link degradation before tenants do (e.g. Prometheus/Grafana/Datadog, streaming telemetry)
  • Troubleshoot performance issues end to end, from optics and cabling through switch buffers to NIC/DPU configuration and collective-communication behaviour on the hosts
  • Operate the out-of-band management network, console access, and remote recovery paths
  • Support tenant onboarding: segmentation and addressing, bandwidth and isolation guarantees, and capacity planning as the cluster scales
  • Write clear design documentation capturing decisions, rationale, and rejected alternatives
  • Own and maintain the office network: wired and wireless infrastructure, firewalling, VPN/remote access, and connectivity between the office and data centre environments
  • Upskill colleagues on networking: share knowledge through documentation, run-throughs, and pairing so the wider team can operate and troubleshoot the fabric confidently

Skills

BGP routing
EVPN/VXLAN
Leaf-spine design
Linux networking
Automation
RDMA fabric
Ansible
Python
Telemetry

Tools

NetBox
Prometheus
Grafana
InfiniBand

Job description

Fuse Energy is building a best-in-class network fabric for a multi-tenant AI cluster, spanning high-performance compute, storage, and management networks. You will own the fabric architecture through day-2 operations, including RDMA traffic and tenant isolation.

You’ll design, deploy, and operate leaf-spine data centre fabrics, automate provisioning, and build observability dashboards to detect congestion before tenants notice.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Network Engineer
HPC Network Engineer

Multicoin • London

On-site
NOK 1,141,000 - 1,648,000
Competitive salary and equity sign-on
Biannual bonus scheme
Fully expensed tech to match your need
+1
Datacentre Operations Engineer — AI/HPC, High-Density GPU
Datacentre Operations Engineer — AI/HPC, High-Density GPU

Radiant • London

On-site
NOK 887,000 - 1,395,000
GPU HPC Data Center Ops Lead
GPU HPC Data Center Ops Lead

Nscale • Nordland fylke

On-site
NOK 900,000 - 1,300,000
Senior Infrastructure Operations Lead - AI Data Center
Senior Infrastructure Operations Lead - AI Data Center

Nscale • Glomfjord

On-site
NOK 1,200,000 - 1,600,000
Data Center Ops Lead - HPC/GPU Infrastructure
Data Center Ops Lead - HPC/GPU Infrastructure

Nscale • Norway

On-site
NOK 950,000 - 1,300,000
Data Center Technician — HPC & Enterprise Hardware
Data Center Technician — HPC & Enterprise Hardware

LEFDAL MINE DATA CENTERS AS • Kjølsdalen

On-site
NOK 550,000 - 750,000
Modern HPC infrastructure
Collaborative technical environment
Professional development
+1
Datacentre Operations Engineer
Datacentre Operations Engineer

Radiant • London

On-site
NOK 887,000 - 1,395,000
Data Center Design Lead for AI Infrastructure
Data Center Design Lead for AI Infrastructure

Radiant • London

On-site
NOK 1,141,000 - 1,521,000
Infrastructure Operations Lead - AI Data Center (Onsite)
Infrastructure Operations Lead - AI Data Center (Onsite)

Nscale • Ørnes - Gárrgonjárrga

On-site
NOK 1,000,000 - 1,600,000
Linux Systems Engineer for HPC & On-Prem Clusters
Linux Systems Engineer for HPC & On-Prem Clusters

Fractile • London

On-site
NOK 761,000 - 1,141,000
Equity & Ownership
Private Medical
Dental and Vision
+2