Staff Network Engineer — GPU Data Center & HPC Networking

Matcha

Northern (KY)

Hybrid

USD 150,000 - 300,000

Full time

6 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Prime Intellect is seeking a Senior Network Engineer to design and operate the networks that connect large GPU clusters. You will own the reliability and performance of training fabrics, storage networks, and management connectivity to ensure distributed workloads scale without bottlenecks.

You will build scalable datacenter topologies, configure Ethernet/RoCE and InfiniBand fabrics, automate provisioning, diagnose performance issues, and partner with operators and vendors to resolve failures.

Qualifications

  • 3+ years of production datacenter networking experience.
  • Strong understanding of Ethernet, TCP/IP, routing, switching, and redundant network design.
  • Hands-on experience with high-performance GPU networking using InfiniBand or RoCE.
  • Experience troubleshooting network problems across Linux hosts, NICs, switches, and physical links.
  • Ability to automate network operations with Python, Ansible, or comparable tools.

Responsibilities

  • Design and deploy scalable datacenter network topologies for GPU training, inference, storage, and management traffic.
  • Configure and operate high-performance Ethernet/RoCE and InfiniBand fabrics with standards for routing, redundancy, and capacity.
  • Automate network provisioning, configuration validation, upgrades, and rollback procedures.
  • Diagnose packet loss, congestion, link failures, and performance across hosts and switches.
  • Benchmark end-to-end network performance with infra and ML teams; translate workload needs into acceptance criteria.
  • Build monitoring for port health, errors, utilization, congestion, and topology; improve incident response.
  • Partner with datacenter operators and hardware vendors on cabling, optics, deployment readiness, and failure resolution.

Skills

Datacenter networking
Ethernet/TCP/IP
RoCE/InfiniBand
Linux networking
Automation (Python/Ansible)

Tools

Python
Ansible
NIC drivers
RoCE/InfiniBand tooling

Job description

Prime Intellect is seeking a Senior Network Engineer to design and operate the networks that connect large GPU clusters. You will own the reliability and performance of training fabrics, storage networks, and management connectivity to ensure distributed workloads scale without bottlenecks.

You will build scalable datacenter topologies, configure Ethernet/RoCE and InfiniBand fabrics, automate provisioning, diagnose performance issues, and partner with operators and vendors to resolve failures.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Datacenter Networking Engineer for GPU Infra
Staff Datacenter Networking Engineer for GPU Infra

Prime Intellect • United States

On-site
USD 150,000 - 300,000
Senior Network Architect – HPC GPU Data Center
Senior Network Architect – HPC GPU Data Center

Advanced Micro Devices • San Jose (CA)

Hybrid
USD 150,000 - 190,000
Senior Network Architect for 10k+ GPU HPC Clusters
Senior Network Architect for 10k+ GPU HPC Clusters

AMD • San Jose (CA)

Hybrid
USD 150,000 - 190,000
Hybrid work model
AMD benefits
Senior Data Center Network Engineer for HPC and GPU
Senior Data Center Network Engineer for HPC and GPU

GTN Technical Staffing • Dallas (TX)

On-site
USD 120,000 - 180,000
Relocation assistance
GPU Network Architect for AI/HPC Clusters
GPU Network Architect for AI/HPC Clusters

CyberCoders • Santa Clara (CA)

Hybrid
USD 200,000 - 250,000
Health Benefits
401k
Relocation assistance
Staff Datacenter Networking Engineer (GPU & AI Infra)
Staff Datacenter Networking Engineer (GPU & AI Infra)

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Staff Datacenter Networking Engineer: Frontier AI GPU
Staff Datacenter Networking Engineer: Frontier AI GPU

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Member of Technical Staff - Datacenter Networking at Prime Intellect
Member of Technical Staff - Datacenter Networking at Prime Intellect

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000