Senior Kubernetes Networking Engineer for GPU Clusters

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 184,000 - 357,000

Full time

8 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA in Santa Clara, CA seeks a senior networking engineer to own how our clusters communicate, including the CNI data plane, overlay mesh, and regional connections. You will design, operate, and scale Kubernetes networking for GPU clusters at multi-thousand-node scale using modern topologies.

You will design, operate and scale L7 gateways, VPN topologies (Tailscale/WireGuard), and gateways that connect control and data planes, while diagnosing advanced scale challenges and collaborating

Qualifications

  • BS/MS in Computer Science, Electrical Engineering or related field required.
  • 6+ years of professional experience in systems, network or infrastructure software engineering.
  • Deep command of Kubernetes networking and CNI standards; production experience with Calico preferred.
  • Proficiency designing and maintaining modern mesh and VPN topologies (Tailscale, WireGuard).
  • Strong Linux networking fundamentals: routing, iptables/nftables, namespaces, and interaction with container runtimes.
  • Ability to debug distributed network problems at scale using packet capture and tracing.
  • Proficiency in Go, Python, or C/C++.
  • Clear written and verbal communication; collaboration across time zones.

Responsibilities

  • Own how clusters communicate internally and externally, including CNI data plane, overlay mesh, and inter-region connections.
  • Own and evolve the Kubernetes networking architecture for GPU clusters at multi-thousand-node scale.
  • Design, operate and scale overlay networks (CNI, mesh, VPN topologies) and gateways that connect control and data planes.
  • Design, operate and scale L7 gateways/load balancers/tunnels (Envoy, Cloudflare).
  • Identify scale ceilings: packet loss under load, control-plane saturation, IP exhaustion, and failure modes at scale.
  • Build scale-test environments and validation suites to catch networking regressions before production.
  • Diagnose hard, ambiguous problems across the stack (Slurm, training jobs, routing issues).
  • Partner with cloud/neocloud providers on network topology as new clusters are brought up.
  • Provide senior technical judgement to a distributed team with depth in the Kubernetes networking domain.

Skills

Kubernetes networking
CNI & Calico
Mesh & VPN topologies
Envoy / L7 load balancers
Go / Python / C
Linux networking fundamentals
Distributed tracing
Communication

Education

BS/MS in Computer Science/EE or related field

Tools

Calico
Tailscale
WireGuard
Envoy
Cloudflare

Job description

NVIDIA in Santa Clara, CA seeks a senior networking engineer to own how our clusters communicate, including the CNI data plane, overlay mesh, and regional connections. You will design, operate, and scale Kubernetes networking for GPU clusters at multi-thousand-node scale using modern topologies.

You will design, operate and scale L7 gateways, VPN topologies (Tailscale/WireGuard), and gateways that connect control and data planes, while diagnosing advanced scale challenges and collaborating

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Cluster Networking Engineer
Senior GPU Cluster Networking Engineer

Socket.dev • North Carolina

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Cluster Networking Architect for GPU Superclusters
Senior Cluster Networking Architect for GPU Superclusters

NVIDIA • Durham (NC)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Cluster Networking Architect for GPU AI HPC
Senior Cluster Networking Architect for GPU AI HPC

NVIDIA • Austin (TX)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior GPU Cluster Networking Architect — Multi-Cloud Scale
Senior GPU Cluster Networking Architect — Multi-Cloud Scale

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior Cluster Networking Engineer for GPU AI Superclusters
Senior Cluster Networking Engineer for GPU AI Superclusters

NVIDIA • Westford (MA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior GPU Cluster Networking Architect
Senior GPU Cluster Networking Architect

NVIDIA • United States

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Software Engineer - Cluster Networking
Senior Software Engineer - Cluster Networking

NVIDIA AI • Durham (NC)

On-site
USD 180,000 - 240,000
Equity
Benefits
Senior Software Engineer - Cluster Networking
Senior Software Engineer - Cluster Networking

NVIDIA • Durham (NC)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Software Engineer - Cluster Networking
Senior Software Engineer - Cluster Networking

NVIDIA • Westford (MA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Kubernetes Engineer for GPU AI Compute Platform
Senior Kubernetes Engineer for GPU AI Compute Platform

GTN Technical Staffing • Dallas (TX)

On-site
USD 140,000 - 200,000