Senior Cluster Networking Engineer for GPU AI Superclusters

NVIDIA

Westford (MA)

On-site

USD 184,000 - 357,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a senior networking engineer to lead the network architecture for GPU superclusters. The platform trains across tens of thousands of GPUs on multiple clouds, and we are scaling to clusters of ten thousand nodes and beyond.

You will own Kubernetes networking, design overlay networks and gateways, and build scale-test environments to catch regressions before production. Collaboration across time zones and cloud providers is essential.

Qualifications

  • BS/MS in Computer Science, Electrical Engineering or a related field, or equivalent experience.
  • 6+ years of professional experience in systems, network or infrastructure software engineering.
  • Deep command of Kubernetes networking architecture and CNI standards, with production experience operating Calico strongly preferred.
  • Proficiency designing and maintaining modern mesh and VPN networking topologies - Tailscale, WireGuard or equivalent.
  • Strong Linux networking fundamentals: routing, netfilter and iptables/nftables, packet marking, network namespaces, and how these interact with container runtimes.
  • Demonstrated ability to debug distributed network problems at scale - packet capture, tracing, and correlating behaviour across many hosts to find a single root cause.
  • Proficiency in Go, Python, C or a comparable systems language.
  • Clear written and verbal communication, and the ability to work effectively with engineers across multiple time zones.

Responsibilities

  • Own and evolve the Kubernetes networking architecture for GPU clusters running at multi-thousand-node scale.
  • Design, operate and scale the overlay network - CNI, mesh and VPN topologies (Tailscale, WireGuard), and the gateways that connect control and data planes.
  • Design, operate and scale the L7 gateways/load balancers/tunnels (Envoy, Cloudflare).
  • Find and eliminate scale ceilings: packet loss under load, control-plane saturation, IP address management exhaustion, and the failure modes that only appear above a few thousand nodes.
  • Build the scale-test environments and validation suites that let us catch networking regressions before they reach production, rather than during a training run.
  • Diagnose hard, ambiguous problems across the stack - where a symptom in Slurm or a training job traces back to a mark collision, a stale route, or a saturated tunnel.
  • Partner with cloud and neocloud providers on network topology, requirements and capabilities as we bring up new clusters.
  • Provide senior technical judgement to a distributed team, and depth in the Custer Networking domain.

Skills

Kubernetes networking
CNI (Calico)
Linux networking
Go
Python
C
Packet tracing

Education

BS/MS in Computer Science or Electrical Engineering

Tools

Tailscale
WireGuard
Envoy
Cloudflare

Job description

NVIDIA is seeking a senior networking engineer to lead the network architecture for GPU superclusters. The platform trains across tens of thousands of GPUs on multiple clouds, and we are scaling to clusters of ten thousand nodes and beyond.

You will own Kubernetes networking, design overlay networks and gateways, and build scale-test environments to catch regressions before production. Collaboration across time zones and cloud providers is essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cluster Networking Architect for GPU Superclusters
Senior Cluster Networking Architect for GPU Superclusters

NVIDIA • Durham (NC)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior GPU Cluster Networking Architect — Multi-Cloud Scale
Senior GPU Cluster Networking Architect — Multi-Cloud Scale

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior GPU Cluster Networking Engineer
Senior GPU Cluster Networking Engineer

Socket.dev • North Carolina

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Cluster Networking Architect for GPU AI HPC
Senior Cluster Networking Architect for GPU AI HPC

NVIDIA • Austin (TX)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior GPU Cluster Networking Architect
Senior GPU Cluster Networking Architect

NVIDIA • United States

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Kubernetes Networking Engineer for GPU Clusters
Senior Kubernetes Networking Engineer for GPU Clusters

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Solutions Architect - AI Cluster Networking Design
Senior Solutions Architect - AI Cluster Networking Design

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Software Engineer - Cluster Networking
Senior Software Engineer - Cluster Networking

NVIDIA • Durham (NC)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Software Engineer - Cluster Networking
Senior Software Engineer - Cluster Networking

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Software Engineer - Cluster Networking
Senior Software Engineer - Cluster Networking

Socket.dev • North Carolina

On-site
USD 184,000 - 357,000
Equity
Benefits