GPU Network Engineer

Blue Signal Search

Santa Clara (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Blue Signal Search is seeking a GPU Network Engineer to design and operate the high-performance networking backbone powering large AI training and inference workloads. You will work with experienced infra and compute teams to influence GPU cluster architecture and scalable data center fabrics.

The role emphasizes optimizing ultra-low latency communication, implementing modern data center fabrics, and automating provisioning and lifecycle management.

Qualifications

  • Minimum of 5 years of experience supporting enterprise data center or HPC networking environments with direct GPU cluster networking experience.
  • Hands-on experience with InfiniBand (NDR or XDR) and RoCEv2 networking technologies.
  • Strong expertise designing and supporting east west networking for GPU or AI infrastructure.
  • Experience with CLOS architectures, ECMP routing, EVPN, and VXLAN.
  • Hands-on administration of Arista EOS, Juniper Junos, and NVIDIA or Mellanox networking platforms.
  • Experience using NetBox, Netconf, Ansible, and Terraform to automate network deployment and management.
  • Strong troubleshooting skills, packet analysis experience, and network performance optimization capabilities.
  • Excellent communication and technical documentation skills.

Responsibilities

  • Design, deploy, and optimize high performance InfiniBand (NDR and XDR) and RoCEv2 fabrics supporting large scale GPU and AI workloads.
  • Build scalable CLOS and ECMP network architectures that deliver low latency, high bandwidth east west traffic across distributed GPU environments.
  • Implement and support EVPN, VXLAN, Layer 2, and Layer 3 networking for modern data center infrastructure.
  • Configure and maintain Arista, Juniper, and NVIDIA networking platforms while ensuring highly available, resilient network operations.
  • Automate network provisioning, configuration, and lifecycle management using Netconf, Ansible, Terraform, and NetBox.
  • Monitor, troubleshoot, and optimize network performance through packet analysis, telemetry, and root cause analysis.
  • Partner with infrastructure, compute, and platform engineering teams to support AI training, inference, and high performance computing initiatives.
  • Create and maintain technical documentation, operational procedures, and implementation standards while participating in change management and infrastructure improvements.

Tools

Arista EOS
Juniper Junos
NVIDIA/Mellanox platforms

Job description

An innovative technology organization at the forefront of AI infrastructure is seeking a GPU Network Engineer to help design and operate the high-performance networking backbone powering advanced GPU computing environments. This opportunity is ideal for an engineer who enjoys solving complex networking challenges, optimizing ultra low latency communication, and building highly scalable infrastructure supporting large AI training and inference workloads. You will work alongside experienced infrastructure and compute professionals while influencing the architecture of next generation GPU clusters.

What You Will Do
  • Design, deploy, and optimize high performance InfiniBand (NDR and XDR) and RoCEv2 fabrics supporting large scale GPU and AI workloads.
  • Build scalable CLOS and ECMP network architectures that deliver low latency, high bandwidth east west traffic across distributed GPU environments.
  • Implement and support EVPN, VXLAN, Layer 2, and Layer 3 networking for modern data center infrastructure.
  • Configure and maintain Arista, Juniper, and NVIDIA networking platforms while ensuring highly available, resilient network operations.
  • Automate network provisioning, configuration, and lifecycle management using Netconf, Ansible, Terraform, and NetBox.
  • Monitor, troubleshoot, and optimize network performance through packet analysis, telemetry, and root cause analysis.
  • Partner with infrastructure, compute, and platform engineering teams to support AI training, inference, and high performance computing initiatives.
  • Create and maintain technical documentation, operational procedures, and implementation standards while participating in change management and infrastructure improvements.
Required Qualifications
  • Minimum of 5 years of experience supporting enterprise data center or HPC networking environments with direct GPU cluster networking experience.
  • Extensive hands on experience with InfiniBand (NDR or XDR) and RoCEv2 networking technologies.
  • Strong expertise designing and supporting east west networking for GPU or AI infrastructure.
  • Experience with CLOS architectures, ECMP routing, EVPN, and VXLAN.
  • Hands on administration of Arista EOS, Juniper Junos, and NVIDIA or Mellanox networking platforms.
  • Experience using NetBox, Netconf, Ansible, and Terraform to automate network deployment and management.
  • Strong troubleshooting skills, packet analysis experience, and network performance optimization capabilities.
  • Excellent communication and technical documentation skills.
Preferred Qualifications
  • Experience optimizing fabrics supporting large scale AI training workloads.
  • Knowledge of adaptive routing, congestion management, RDMA optimization, and lossless networking.
  • Python scripting for infrastructure automation.
  • Industry certifications such as CCNP, CCIE, JNCIP, JNCIE, or equivalent practical experience.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior Network Solution Architect – AI Fabrics
Senior Network Solution Architect – AI Fabrics

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 180,000 - 240,000
Network Engineer
Network Engineer

QumulusAI • United States

On-site
USD 100,000 - 130,000
Competitive compensation
Equity participation
Opportunity to work on cutting-edge technology
Senior AI Network Engineer - AI Infrastructure
Senior AI Network Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • United States

On-site
USD 220,000 - 350,000
Annual bonus
Equity opportunities
Flexible working arrangements
+1
Network Engineer
Network Engineer

Sesterce Group • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior GPU AI Network Architect
Senior GPU AI Network Architect

Sesterce Group • San Francisco (CA)

On-site
USD 120,000 - 160,000
Network Architect
Network Architect

TechDigital Group • California (MO)

Hybrid
USD 120,000 - 160,000
Network Engineer – Design & Engineering
Network Engineer – Design & Engineering

Jobtailor • California (MO)

On-site
USD 150,000 - 230,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Network Solutions Architect, AI Factory Services
Network Solutions Architect, AI Factory Services

Lenovo • North Carolina

Hybrid
USD 100,000 - 130,000
Hybrid work schedule
Equal Opportunity Employer