Staff Datacenter Networking Engineer for GPU Infra

Prime Intellect

United States

On-site

USD 150,000 - 300,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Prime Intellect is building the open superintelligence stack and unified Lab platform for frontier AI workloads. You will design and operate the networks that connect large GPU clusters, ensuring reliability and performance for training and deployment.

The role requires hands-on experience with InfiniBand/RoCE, Ethernet/TCP-IP, and automation with Python or Ansible, plus a track record of large-scale deployments. Collaboration with datacenter operators and engineering teams is essential.

Qualifications

  • 3+ years in production datacenter networking.
  • Strong Ethernet, TCP/IP, routing, switching knowledge.
  • Hands-on GPU networking with InfiniBand or RoCE.
  • Automate network operations with Python/Ansible.
  • Ability to diagnose network issues on Linux hosts and switches.

Responsibilities

  • Design and deploy scalable datacenter network topologies for GPU training, inference, storage, and management traffic.
  • Configure and operate high-performance Ethernet/RoCE and InfiniBand fabrics with routing, redundancy, and capacity standards.
  • Automate network provisioning, validation, upgrades, and rollback procedures.
  • Diagnose packet loss, congestion, link failures, and performance issues across hosts and switches.
  • Benchmark end-to-end network performance with infrastructure and ML teams; translate workloads into acceptance criteria.
  • Build monitoring for port health, errors, utilization, congestion, and topology; improve incident response and runbooks.
  • Partner with datacenter operators and hardware vendors on cabling, optics, deployment readiness, and failure resolution.

Skills

Datacenter networking
GPU networking
TCP/IP networking
Python automation
Network design
Troubleshooting

Tools

InfiniBand
RoCE
BGP/ECMP
NIC drivers

Job description

Prime Intellect is building the open superintelligence stack and unified Lab platform for frontier AI workloads. You will design and operate the networks that connect large GPU clusters, ensuring reliability and performance for training and deployment.

The role requires hands-on experience with InfiniBand/RoCE, Ethernet/TCP-IP, and automation with Python or Ansible, plus a track record of large-scale deployments. Collaboration with datacenter operators and engineering teams is essential.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Datacenter Networking Engineer: Frontier AI GPU
Staff Datacenter Networking Engineer: Frontier AI GPU

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Staff Network Engineer — GPU Data Center & HPC Networking
Staff Network Engineer — GPU Data Center & HPC Networking

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000
GPU Cloud Infrastructure Engineer
GPU Cloud Infrastructure Engineer

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Networking at Prime Intellect
Member of Technical Staff - Datacenter Networking at Prime Intellect

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Networking
Member of Technical Staff - Datacenter Networking

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Staff Datacenter Networking Engineer (GPU & AI Infra)
Staff Datacenter Networking Engineer (GPU & AI Infra)

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Networking
Member of Technical Staff - Datacenter Networking

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Networking
Member of Technical Staff - Datacenter Networking

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior GPU Cloud Operations Engineer
Senior GPU Cloud Operations Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000