Senior GPU Cluster Networking Engineer

Socket.dev

North Carolina

On-site

USD 184,000 - 357,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a senior networking engineer to own the network architecture for its GPU superclusters. You will lead Kubernetes networking for multi-thousand-node clusters, design and operate overlay networks (CNI, mesh, VPNs), and scale L7 gateways and tunnels using Envoy and Cloudflare.

You will diagnose complex scale issues and collaborate with cloud providers to expand clusters. The role requires deep Kubernetes networking expertise, strong Linux networking fundamentals, and proficiency

Qualifications

  • BS/MS in CS, Electrical Engineering or related field, or equivalent experience.
  • 6+ years of professional experience in systems, network or infrastructure software engineering.
  • Deep command of Kubernetes networking architecture and CNI standards, with production experience operating Calico strongly preferred.
  • Proficiency designing and maintaining modern mesh and VPN networking topologies - Tailscale, WireGuard or equivalent.
  • Strong Linux networking fundamentals: routing, netfilter and iptables/nftables, packet marking, network namespaces, and how these interact with container runtimes.
  • Demonstrated ability to debug distributed network problems at scale - packet capture, tracing, and correlating behaviour across many hosts to find a single root cause.
  • Proficiency in Go, Python, C or a comparable systems language.
  • Clear written and verbal communication, and the ability to work effectively with engineers across multiple time zones.

Responsibilities

  • Own and evolve the Kubernetes networking architecture for GPU clusters running at multi-thousand-node scale.
  • Design, operate and scale the overlay network - CNI, mesh and VPN topologies, and the gateways that connect control and data planes.
  • Design, operate and scale the L7 gateways/load balancers/tunnels (Envoy, Cloudflare).
  • Find and eliminate scale ceilings: packet loss under load, control-plane saturation, IP address management exhaustion, and the failure modes that only appear above a few thousand nodes.
  • Build the scale-test environments and validation suites that let us catch networking regressions before they reach production.
  • Diagnose hard, ambiguous problems across the stack - trace issues across Slurm or training jobs.

Skills

Kubernetes networking
Linux networking
Go
Python
C
Network debugging

Education

BS in CS/EE
MS in CS/EE

Tools

Calico
Tailscale
WireGuard
Envoy
Cloudflare

Job description

NVIDIA is seeking a senior networking engineer to own the network architecture for its GPU superclusters. You will lead Kubernetes networking for multi-thousand-node clusters, design and operate overlay networks (CNI, mesh, VPNs), and scale L7 gateways and tunnels using Envoy and Cloudflare.

You will diagnose complex scale issues and collaborate with cloud providers to expand clusters. The role requires deep Kubernetes networking expertise, strong Linux networking fundamentals, and proficiency

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Cluster Networking Architect — Multi-Cloud Scale
Senior GPU Cluster Networking Architect — Multi-Cloud Scale

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Cluster Networking Architect for GPU Superclusters
Senior Cluster Networking Architect for GPU Superclusters

NVIDIA • Durham (NC)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Cluster Networking Engineer for GPU AI Superclusters
Senior Cluster Networking Engineer for GPU AI Superclusters

NVIDIA • Westford (MA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior GPU Cluster Networking Architect
Senior GPU Cluster Networking Architect

NVIDIA • United States

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Kubernetes Networking Engineer for GPU Clusters
Senior Kubernetes Networking Engineer for GPU Clusters

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Cluster Networking Architect for GPU AI HPC
Senior Cluster Networking Architect for GPU AI HPC

NVIDIA • Austin (TX)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Software Engineer - Cluster Networking
Senior Software Engineer - Cluster Networking

NVIDIA AI • Durham (NC)

On-site
USD 180,000 - 240,000
Equity
Benefits
Senior Software Engineer - Cluster Networking
Senior Software Engineer - Cluster Networking

NVIDIA • Durham (NC)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Software Engineer - Cluster Networking
Senior Software Engineer - Cluster Networking

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Software Engineer - Cluster Networking
Senior Software Engineer - Cluster Networking

Socket.dev • North Carolina

On-site
USD 184,000 - 357,000
Equity
Benefits