Location:
Remote (US)
Base:
$160,000-$190,000 + Bonus + Equity (RSUs - 4-year vest)
The Client
I'm working with a neocloud building purpose-built GPU infrastructure for AI workloads. They run large-scale clusters powering training and inference for some of the most demanding AI customers in the market.
The role
They're hiring a Network Engineer to design, build, and operate the network fabric behind their GPU cloud. This is a hands-on engineering role at the centre of a fast-moving buildout: multi-site data centres, new NVIDIA and AMD GPU clusters, and a greenfield cloud platform with automation-first network design.
You'll work across the whole stack from Juniper edge routing and firewalling, to leaf-spine fabrics on NVIDIA/Mellanox Cumulus and Arista, to the RDMA fabrics (InfiniBand and RoCE) that make large-scale training and inference possible. You'll spend as much time in automation and code as on the CLI, and you'll make trips to the data centres to help bring new capacity online.
What you'll own
- Design, configuration, deployment, and tuning of high-speed network infrastructure for GPU training and inference across our clusters
- The GPU fabric itself: RDMA over InfiniBand and RoCEv2 - congestion control, PFC/ECN tuning, and Pkey/partitioning for multi-tenant isolation
- Edge routing, firewalling, and multi-homed connectivity on Juniper: BGP, public IP space, and upstream carrier relationships
- Leaf-spine data centre fabrics on NVIDIA/Mellanox Cumulus Linux, and Arista
- Network automation and infrastructure-as-code: codifying device configs, fabric buildout, and validation in Python, Ansible, and Terraform rather than keeping it in people's heads
- Observability and troubleshooting across Prometheus, Grafana, Checkmk, Wireshark, and tcpdump, with a clear point of view on what good monitoring looks like
- Network security posture alongside our security engineers: firewall policy, IDS/IPS, WireGuard VPN, and controls that hold up to SOC 2 and HIPAA
- Data centre design: structured fibre and copper cabling, and physical-layer troubleshooting during new site and cluster buildouts
What they're looking for
Required
- Significant experience designing, operating, and troubleshooting production network infrastructure at scale
- Deep understanding of TCP/IP, BGP, VLANs, VXLAN/EVPN, and modern data centre fabric design
- Hands-on high-performance networking for HPC or AI/ML: RDMA, InfiniBand and/or RoCE, and the tuning that makes them perform
- Production experience with NVIDIA/Mellanox Cumulus Linux and/or Arista EOS in leaf-spine topologies
- Strong Juniper routing and firewall experience (Junos)
- Fluency with network automation and IaC: Python, Ansible, Terraform, and Bash
- Comfortable with packet-level troubleshooting (Wireshark, tcpdump) and metrics-based troubleshooting (Prometheus, Grafana)
Non-essential but nice to have:
- Building or operating GPU/HPC network fabrics at scale (ConnectX, Bluefield DPU, Spectrum, Quantum, UFM, or equivalent)
- Familiarity with SDN and network-automation platforms (e.g. Netris or comparable)
- OpenStack/OVN networking, or standing up SDN on a greenfield cloud platform
- Working knowledge of NVIDIA GPU platforms (H100, B200) and their InfiniBand scale-out networking
- Exposure to AMD GPU platforms (MI-series) and their Ethernet scale-out networking
- Zero-trust network design and WireGuard VPN
- Operating under SOC 2 and/or HIPAA
- MAAS or bare-metal provisioning