Remote Senior GPU Network Engineer – Fabric & Cluster Builds

REALM

United States

On-site

USD 170,000 - 230,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Realm is hiring a senior network architect to own the GPU training fabric at scale. You’ll design buildouts, lead deployments, and act as the main escalation point for performance incidents, enabling high-throughput AI workloads.

You’ll collaborate with global customers, automate repetitive tasks with Python and Ansible, and mentor teammates. This is a fully remote role within the United States with flexible scheduling for maintenance windows.

Qualifications

  • 7+ years in network operations or engineering with growing responsibilities.
  • Hands-on GPU cluster networking: RDMA/RoCEv2 fabric config and lossless Ethernet design.
  • Expert-level BGP on internet-scale networks and ECMP awareness.
  • Strong Linux networking know-how: kernel routing, nftables/iptables, FRR/Bird.
  • Comfort with multiple vendors/NOS: Juniper, Cisco, Nokia, Cumulus, SONiC.
  • Proficiency in Python and Ansible for repeatable changes and data collection.

Responsibilities

  • Own the fabric behind GPU clusters; diagnose latency issues and fix root causes.
  • Lead design and deployment of GPU cluster buildouts and migrations globally.
  • Serve as senior escalation for critical network incidents with cross-functional coordination.
  • Onboard strategic customers to GPU deployments and resolve performance problems.
  • Champion automation; develop tooling to reduce manual toil.
  • Mentor engineers and raise the team's baseline through docs and pairing.
  • Participate in 24x7x365 on-call rotation and incident leads.

Skills

Networking
GPU networking
BGP
Linux networking
Python
Ansible
Automation

Education

Bachelor's degree in a relevant field

Tools

Juniper
Cisco
Nokia
Cumulus
SONiC
nftables
iptables
NVLink/NVSwitch
RDMA/RoCEv2

Job description

Realm is hiring a senior network architect to own the GPU training fabric at scale. You’ll design buildouts, lead deployments, and act as the main escalation point for performance incidents, enabling high-throughput AI workloads.

You’ll collaborate with global customers, automate repetitive tasks with Python and Ansible, and mentor teammates. This is a fully remote role within the United States with flexible scheduling for maintenance windows.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Senior Network Engineer — AI Data Center Fabrics
Remote Senior Network Engineer — AI Data Center Fabrics

your Jared • Portland (OR), Northern (KY)

Hybrid
USD 83,000 - 152,000
Medical coverage
Dental coverage
Vision coverage
+2
Staff Network Platform Engineer – GPU Cloud Automation
Staff Network Platform Engineer – GPU Cloud Automation

CoreWeave • New York (NY)

On-site
USD 180,000 - 260,000
Medical insurance
Dental insurance
Vision insurance
+13
Senior GPU Infra Engineer — Customer-Facing (Hybrid/Remote)
Senior GPU Infra Engineer — Customer-Facing (Hybrid/Remote)

Rune • Mountain View (CA)

Hybrid
USD 175,000 - 260,000
Senior GPU Infra Engineer — Remote
Senior GPU Infra Engineer — Remote

Nscale • Seattle (WA)

On-site
USD 120,000 - 170,000
Remote-first culture
Equity plan
Flexible workplace
Senior GPU Infrastructure Support Engineer
Senior GPU Infrastructure Support Engineer

Nscale • San Francisco (CA)

On-site
USD 120,000 - 170,000
Equity
Remote-friendly team
Flexible workplace
Senior Network Architect for 10k+ GPU HPC Clusters
Senior Network Architect for 10k+ GPU HPC Clusters

AMD • San Jose (CA)

Hybrid
USD 150,000 - 190,000
Hybrid work model
AMD benefits
Senior GPU Compute Fleet Engineer | Remote-First
Senior GPU Compute Fleet Engineer | Remote-First

Boundless Networks • United States

Remote
USD 175,000 - 250,000
Equity
Health, dental, vision
Flexible PTO
+2
Remote GPU Infra NOC Engineer — Automation & AI Ops
Remote GPU Infra NOC Engineer — Automation & AI Ops

Orionplacement • Pittsburgh

On-site
USD 75,000 - 140,000
Bonus and equity opportunities
Medical, dental, and vision insurance
401(k)
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)

Sciforium • San Francisco (CA)

On-site
USD 170,000 - 230,000