HPC Network Engineer

Fuse Energy

Greater London

On-site

GBP 90,000 - 120,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive salary and an equity sign‑
Biannual bonus scheme
Fully expensed tech to match your need
Breakfast and dinner allowance for off

Job summary

Fuse Energy is seeking a senior network engineer to design, deploy and operate the fabric for a multi-tenant AI cluster and data centre. You will own architecture, day-2 operations, and keep the network secure, scalable and observable.

You will implement lossless RDMA fabrics (RoCEv2/InfiniBand), leaf-spine designs, automation with Ansible and NetBox, and deliver clear design documentation while mentoring colleagues across teams.

Qualifications

  • 5+ years of production data centre networking experience.
  • Strong BGP/EVPN/VXLAN design and troubleshooting.
  • Hands-on experience with leaf-spine/Clos fabrics and Linux networking.

Responsibilities

  • Design and operate lossless, RDMA-capable fabrics for GPU compute and storage.
  • Build leaf-spine data centre fabrics with routed underlay and overlay design.
  • Implement per-tenant network isolation across compute, storage, and management.
  • Automate provisioning, configuration, and validation; treat switch config as code.
  • Build telemetry and observability for fabric with dashboards and alerts.
  • Troubleshoot performance end-to-end from optics to NIC/DPU configuration.
  • Operate the out-of-band management network and remote recovery paths.
  • Support tenant onboarding: segmentation, addressing, bandwidth guarantees.
  • Write design documentation detailing decisions and reasoning.
  • Own office network: wired/wireless, firewalling, VPN, and connectivity to data centre.
  • Upskill colleagues through documentation, run-throughs, and pairing.

Skills

Network engineering
BGP
EVPN/VXLAN
Leaf-spine design
Linux networking
Python scripting
Automation
Documentation
Telemetry/Monitoring
Campus networking

Tools

Ansible
NetBox
Prometheus/Grafana
MAAS
PXE

Job description

Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy - fast. We're combining first-principles thinking with cutting-edge technology to build a radically better energy system. We raised $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group and strategic angels like Nico Rosberg, the Co-Founder of Solana and GPs behind Meta, Revolut, Spotify, Uber and more.

The Opportunity

You'll design, deploy, and operate the network fabric for our multi-tenant AI cluster. This covers the full stack: the high-performance compute and storage fabrics carrying RDMA traffic between GPUs, the tenant-facing and management networks, fire walling and tenant isolation, and the out-of-band infrastructure that keeps it all recoverable. Beyond the data centre, you'll own the office network and act as the networking authority for the company, raising the bar for everyone by sharing what you know. You'll own the fabric from architecture through day-2 operations.

Responsibilities


  • Design and operate lossless, RDMA-capable fabrics (e.g. RoCEv2, InfiniBand) for GPU compute and storage traffic, including QoS, congestion control, and buffer tuning at scale

  • Build and manage leaf-spine data centre fabrics, with routed underlay and overlay design (e.g. BGP, EVPN/VXLAN)

  • Implement and maintain per-tenant network isolation across compute, storage, and management planes

  • Automate network provisioning, configuration, and validation, treating switch config as code (e.g. Ansible, Python, NetBox as source of truth), deployed through CI

  • Build telemetry and observability for the fabric: flow-level and buffer-level visibility, dashboards, and alerting that catches congestion and link degradation before tenants do (e.g. Prometheus/Grafana/Datadog, streaming telemetry)

  • Troubleshoot performance issues end to end, from optics and cabling through switch buffers to NIC/DPU configuration and collective-communication behaviour on the hosts

  • Operate the out-of-band management network, console access, and remote recovery paths

  • Support tenant onboarding: segmentation and addressing, bandwidth and isolation guarantees, and capacity planning as the cluster scales

  • Write clear design documentation capturing decisions, rationale, and rejected alternatives

  • Own and maintain the office network: wired and wireless infrastructure, firewalling, VPN/remote access, and connectivity between the office and data centre environments

  • Upskill colleagues on networking: share knowledge through documentation, run-throughs, and pairing so the wider team can operate and troubleshoot the fabric confidently


Requirements


  • 5+ years as a network engineer operating production data centre networks

  • Strong dynamic routing experience (BGP in particular), plus overlay/encapsulation design and troubleshooting (e.g. EVPN/VXLAN)

  • Hands-on experience with leaf-spine / Clos fabric design and operation

  • Experience with modern data centre network operating systems and comfortable in the Linux networking stack, not just a vendor CLI

  • Practical RDMA fabric experience: lossless Ethernet (e.g. RoCEv2 with PFC/ECN/DCQCN tuning) or InfiniBand, with an understanding of why lossless behaviour matters for GPU workloads

  • Network automation as a working practice, not an aspiration: scripting (e.g. Python), configuration management (e.g. Ansible), config generation from a source of truth, version-controlled changes

  • Solid Linux administration fundamentals: you can debug from the host side as well as the switch side

  • Experience with network telemetry and monitoring (e.g. Prometheus/Grafana, sFlow/IPFIX, streaming telemetry)

  • Experience running corporate/campus networks: wired and wireless, switching, NAC/802.1X, VPN and remote access (e.g. Cisco Catalyst/Meraki or comparable)

  • Clear communicator who enjoys teaching: able to document, pair, and run sessions that bring less network-savvy colleagues up to speed


Nice to have


  • Experience with GPU cluster networking specifically (e.g. NVIDIA Spectrum-X or Quantum InfiniBand, ConnectX/BlueField NICs and DPUs, UFM, SHARP, or equivalent Broadcom/Arista AI fabric platforms)

  • Container networking experience: CNI plugins and BGP integration between clusters and the fabric

  • Understanding of collective-communication libraries and how fabric behaviour shows up as training/inference performance

  • Multi-tenant network design: VRF-based isolation, tenant bandwidth guarantees, secure shared infrastructure

  • Experience with enterprise firewall platforms (e.g. FortiGate, Palo Alto), including HA deployment and virtualised/segmented instances

  • Storage networking experience (e.g. NVMe-oF, lossless storage fabrics, per-tenant storage isolation)

  • Bare-metal provisioning environments (e.g. MAAS, PXE, Redfish) and how network bootstrap fits into node lifecycle

  • Optical layer knowledge at 200/400/800G: transceivers, MPO cabling, link qualification

  • Experience standing up a data centre network from greenfield

  • Relevant certifications (e.g. CCNP/CCIE or equivalent), valued as evidence of depth, not a gate


Benefits


  • Competitive salary and an equity sign-on bonus

  • Biannual bonus scheme

  • Fully expensed tech to match your needs

  • Breakfast and dinner allowance for office-based employees


Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network Consultant: HPC, Computing, AI, Artificial Intelligence
Network Consultant: HPC, Computing, AI, Artificial Intelligence

Curo Services • Greater London

On-site
GBP 107,000 - 131,000
Lead Network Operations Engineer
Lead Network Operations Engineer

G-Research • Greater London

On-site
GBP 60,000 - 80,000
Highly competitive compensation plus annual discretionary bonus
Lunch provided
30 days' annual leave
+3
Network Engineer
Network Engineer

NexGen Cloud • Greater London

On-site
GBP 65,000 - 100,000
Competitive salary
Employee wellbeing benefits
25 days holiday
+3
Network Engineer
Network Engineer

asobbi • United Kingdom

Remote
GBP 53,000 - 69,000
Highly competitive package with equity
Dynamic progression plan
Human-first flexibility
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Community Fibre Limited • Greater London

On-site
GBP 70,000 - 110,000
25 days holiday
Flexible WFH policy
Private health cover
+1
HPC Network Engineer - Banking & Finance
HPC Network Engineer - Banking & Finance

Hamilton Barnes Associates Limited • Greater London

On-site
GBP 450,000 - 550,000
Work on advanced low-latency networkch
Influence over architecture and automa
Technical engineering team
+1
Network Engineer
Network Engineer

cgg • Haywards Heath

On-site
GBP 55,000 - 75,000
Competitive salary
Bonus scheme
22 days annual leave
+4
Senior Network Engineer
Senior Network Engineer

Venturi • Manchester

On-site
GBP 70,000 - 100,000
Network Engineer MSP- WAN/LAN
Network Engineer MSP- WAN/LAN

IP-People • Manchester

Hybrid
GBP 30,000 - 35,000
Strong investment in tooling and automation
Collaborative, engineering-led culture
Flexible working arrangements
HPC Network Engineer
HPC Network Engineer

Hamilton Barnes ? • City Of London

Hybrid
GBP 332,000 - 406,000