Senior GPU Cluster Networking Architect

NVIDIA

United States

On-site

USD 184,000 - 357,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a senior networking engineer for its Managed AI Research Superclusters (MARS) to lead network architecture of GPU superclusters. The role spans internal/external cluster communication, CNI data plane, and cross-region connectivity across clouds.

You will own Kubernetes networking for multi-thousand-node GPU clusters, design and operate overlay networks, and scale L7 gateways and tunnels using Envoy and related tech. Strong Linux networking and Go/Python/C skills are required.

Qualifications

  • BS/MS in Computer Science, Electrical Engineering or related field, or equivalent experience.
  • 6+ years of professional experience in systems, network or infrastructure software engineering.
  • Deep command of Kubernetes networking architecture and CNI standards, with production experience operating Calico strongly preferred.
  • Proficiency designing and maintaining modern mesh and VPN networking topologies - Tailscale, WireGuard or equivalent.
  • Strong Linux networking fundamentals: routing, netfilter and iptables/nftables, packet marking, network namespaces, and how these interact with container runtimes.
  • Demonstrated ability to debug distributed network problems at scale - packet capture, tracing, and correlating behaviour across many hosts to find a single root cause.
  • Proficiency in Go, Python, C or a comparable systems language.
  • Clear written and verbal communication, and the ability to work effectively with engineers across multiple time zones.

Responsibilities

  • Be the technical owner of how our clusters communicate internally and externally.
  • Own and evolve the Kubernetes networking architecture for GPU clusters running at multi-thousand-node scale.
  • Design, operate and scale the overlay network - CNI, mesh and VPN topologies, and the gateways that connect control and data planes.
  • Design, operate and scale the L7 gateways/load balancers/tunnels (Envoy, Cloudflare).
  • Find and eliminate scale ceilings: packet loss under load, control-plane saturation, IP address management exhaustion, and failure modes at thousands of nodes.
  • Build scale-test environments and validation suites to catch networking regressions before production.
  • Diagnose hard, ambiguous problems across the stack and trace issues across hosts.

Skills

Kubernetes networking
Go
Python
C
Linux networking

Education

BS/MS in CS/EE or equivalent

Tools

Calico
Tailscale
WireGuard
Envoy

Job description

NVIDIA is seeking a senior networking engineer for its Managed AI Research Superclusters (MARS) to lead network architecture of GPU superclusters. The role spans internal/external cluster communication, CNI data plane, and cross-region connectivity across clouds.

You will own Kubernetes networking for multi-thousand-node GPU clusters, design and operate overlay networks, and scale L7 gateways and tunnels using Envoy and related tech. Strong Linux networking and Go/Python/C skills are required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Cluster Networking Architect — Multi-Cloud Scale
Senior GPU Cluster Networking Architect — Multi-Cloud Scale

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior GPU Cluster Networking Engineer
Senior GPU Cluster Networking Engineer

Socket.dev • North Carolina

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Cluster Networking Engineer for GPU AI Superclusters
Senior Cluster Networking Engineer for GPU AI Superclusters

NVIDIA • Westford (MA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Cluster Networking Architect for GPU Superclusters
Senior Cluster Networking Architect for GPU Superclusters

NVIDIA • Durham (NC)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Cluster Networking Architect for GPU AI HPC
Senior Cluster Networking Architect for GPU AI HPC

NVIDIA • Austin (TX)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Kubernetes Networking Engineer for GPU Clusters
Senior Kubernetes Networking Engineer for GPU Clusters

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Solutions Architect - AI Cluster Networking Design
Senior Solutions Architect - AI Cluster Networking Design

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Software Engineer - Cluster Networking
Senior Software Engineer - Cluster Networking

NVIDIA AI • Durham (NC)

On-site
USD 180,000 - 240,000
Equity
Benefits
Senior Software Engineer - Cluster Networking
Senior Software Engineer - Cluster Networking

Socket.dev • North Carolina

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Software Engineer - Cluster Networking
Senior Software Engineer - Cluster Networking

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000