Senior Datacenter Networking Engineer (GPU Infra)

Primeintellect

San Francisco (CA)

On-site

USD 150,000 - 300,000

Full time

13 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Prime Intellect is building the open superintelligence stack and is seeking a Senior Network Engineer to design and operate the networks connecting large GPU clusters. You will own the reliability and performance of training fabrics, storage networks, and management connectivity so distributed workloads scale without bottlenecks.

Core responsibilities include deploying datacenter topologies for GPU training/inference, configuring Ethernet/RoCE/InfiniBand fabrics, automating provisioning with

Qualifications

  • 3+ years of production datacenter networking experience.
  • Strong understanding of Ethernet, TCP/IP, routing, switching, and redundant network design.
  • Hands-on experience with high-performance GPU networking using InfiniBand or RoCE.
  • Experience troubleshooting network problems across Linux hosts, NICs, switches, and physical links.
  • Ability to automate network operations with Python, Ansible, or comparable tools.

Responsibilities

  • Design and deploy scalable datacenter network topologies for GPU training, inference, storage, and management traffic.
  • Configure and operate high-performance Ethernet/RoCE and InfiniBand fabrics with clear standards for routing, redundancy, and capacity.
  • Automate network provisioning, configuration validation, upgrades, and rollback procedures.
  • Diagnose packet loss, congestion, link failures, and collective communication performance across hosts and switches.
  • Benchmark end-to-end network performance with infrastructure and ML teams, translating workload needs into measurable acceptance criteria.
  • Build monitoring for port health, errors, utilization, congestion, and fabric topology; improve incident response and runbooks.
  • Partner with datacenter operators and hardware vendors on cabling, optics, deployment readiness, and failure resolution.

Skills

Datacenter networking
TCP/IP & routing
GPU networking (InfiniBand/RoCE)
Linux networking
Automation with Python/Ansible

Tools

Python
Ansible

Job description

Prime Intellect is building the open superintelligence stack and is seeking a Senior Network Engineer to design and operate the networks connecting large GPU clusters. You will own the reliability and performance of training fabrics, storage networks, and management connectivity so distributed workloads scale without bottlenecks.

Core responsibilities include deploying datacenter topologies for GPU training/inference, configuring Ethernet/RoCE/InfiniBand fabrics, automating provisioning with

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Network Engineer — GPU Data Center & HPC Networking
Staff Network Engineer — GPU Data Center & HPC Networking

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000
Staff Datacenter Networking Engineer for GPU Infra
Staff Datacenter Networking Engineer for GPU Infra

Prime Intellect • United States

On-site
USD 150,000 - 300,000
Staff Datacenter Networking Engineer: Frontier AI GPU
Staff Datacenter Networking Engineer: Frontier AI GPU

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Staff Datacenter Networking for GPU AI Infrastructure
Staff Datacenter Networking for GPU AI Infrastructure

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior GPU Infrastructure Architect
Senior GPU Infrastructure Architect

Primeintellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
GPU Cloud Infrastructure Engineer
GPU Cloud Infrastructure Engineer

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Staff Datacenter Networking Engineer (GPU & AI Infra)
Staff Datacenter Networking Engineer (GPU & AI Infra)

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Senior GPU Data Center Engineer
Senior GPU Data Center Engineer

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Networking at Prime Intellect
Member of Technical Staff - Datacenter Networking at Prime Intellect

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Networking
Member of Technical Staff - Datacenter Networking

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000