GPU Network Engineer

CyberCoders

Santa Clara (CA)

Hybrid

USD 200,000 - 250,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health Benefits
401k
Relocation assistance

Job summary

CyberCoders in Sunnyvale, CA is seeking a hands-on GPU Network Engineer to design, build and operate the high-speed fabrics that connect our GPU clusters. This hybrid role requires on-site collaboration and a strong focus on reliability and scalability.

You will deploy InfiniBand (NDR/XDR) and RoCEv2 interconnects, design CLOS/ECMP networks, and work with Kubernetes and bare-metal environments to move data efficiently at scale.

Qualifications

  • 2+ years in data center or HPC networking with GPU/AI clusters.
  • Hands-on GPU cluster networking design and operations.
  • Deep experience with InfiniBand (NDR/XDR) and RoCEv2.
  • Experience with Kubernetes and multi-host networking.
  • Familiarity with Netbox, Netconf, and IaC tooling.

Responsibilities

  • Design and deploy east-west network fabrics for GPU-to-GPU, rack-to-rack, and cluster-to-cluster communication.
  • Build and operate InfiniBand (NDR/XDR) and RoCEv2 interconnect fabrics as the primary transport for GPU workloads.
  • Implement CLOS/ECMP architectures for high-bandwidth, low-latency data movement across GPU clusters.

Skills

Data center networking
HPC networking
GPU cluster environments
InfiniBand (NDR/XDR)
RoCEv2 interconnects
Multi-host networking

Tools

Netbox
Netconf
Ansible
Terraform
KVM

Job description

Job Title: GPU Network Engineer
Job Location: Sunnyvale, CA (hybrid)
Job Salary: 200k-250k + Benefits
Requirements: Data center, HPC networking, InfiniBand, Kubernetes

Based in Sunnyvale, CA, we are an innovative energy infrastructure company that develops cutting-edge data centers to drive the expansion of sustainable energy assets.

Due to growth, we are looking for a hands-on GPU Network Engineer to design, build, and operate the high-speed fabrics that connect our GPU clusters.

Must be local to Sunnyvale, CA or willing to relocate for this position. (4 days working on-site)

If you are a GPU Network Engineer, please read on!

What You Will Be Doing:
  1. Design and deploy east-west network fabrics for GPU-to-GPU, rack-to-rack, and cluster-to-cluster communication at scale.
  2. Build and operate InfiniBand (NDR/XDR) and RoCEv2 interconnect fabrics as the primary transport for GPU cluster workloads.
  3. Implement CLOS/ECMP architectures optimized for high-bandwidth, low-latency, lossless data movement across GPU clusters.
Must Have Skills:
  • 2+ years in data center or HPC networking, with direct experience in GPU or AI cluster environments.
  • Hands-on experience with GPU cluster networking: GPU-to-GPU, rack-to-rack, and cluster-to-cluster fabric design and operations.
  • Deep working knowledge of InfiniBand (NDR/XDR) and RoCEv2 as primary GPU cluster interconnects -- this is the core of the role.
  • Experience with multi-host networking for bare metal, KVM, and Kubernetes.
  • Proficiency with Netbox, Netconf, and IaC tools (Ansible, Terraform).
If hired, you will be rewarded with an offer that includes:
  1. 200k-250k salary
  2. Performance Bonus
  3. RSUs
  4. Health Benefits
  5. 401k
  6. PTO
  7. Opportunity to work with cutting-edge AI and HPC infrastructure
  8. Collaborative, fast-paced environment
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
GPU Network Architect for AI/HPC Clusters
GPU Network Architect for AI/HPC Clusters

CyberCoders • Santa Clara (CA)

Hybrid
USD 200,000 - 250,000
Health Benefits
401k
Relocation assistance
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Network Engineer
Network Engineer

asobbi • United States

On-site
USD 160,000 - 190,000
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Network Solutions Architect
Network Solutions Architect

ManpowerGroup Global, Inc. • United States

On-site
USD 120,000 - 180,000
Medical and Prescription Drug Plans
Dental Plan
Vision Plan
+8
GPU Network Engineer: RDMA/NVLink at Scale
GPU Network Engineer: RDMA/NVLink at Scale

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision insurance
Unlimited PTO
Paid parental leave
+1
GPU Infrastructure Engineer
GPU Infrastructure Engineer

Rune • Mountain View (CA)

Hybrid
USD 175,000 - 260,000
GPU Cluster Engineer, Networking
GPU Cluster Engineer, Networking

Sciforium • San Francisco (CA)

On-site
USD 170,000 - 230,000