Staff Datacenter Networking for GPU AI Infrastructure

Prime Intellect AI

San Francisco (CA)

On-site

USD 150,000 - 300,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Prime Intellect AI is building an open frontier AI stack and a unified platform for large-scale GPU training, inference, and tooling. You will design and operate networks that connect massive GPU clusters, ensuring reliability and scalable performance across training fabrics, storage networks, and management connectivity.

You will collaborate with hardware and ML teams to set standards, automate provisioning, and monitor fabric health.

Qualifications

  • 3+ years of production datacenter networking experience.
  • Strong understanding of Ethernet, TCP/IP, routing, switching, and redundant network design.
  • Hands-on experience with high-performance GPU networking using InfiniBand or RoCE.
  • Experience troubleshooting network problems across Linux hosts, NICs, switches, and links.
  • Ability to automate network operations with Python, Ansible, or comparable tools.

Responsibilities

  • Design and deploy scalable datacenter network topologies for GPU training, inference, storage, and management traffic.
  • Configure and operate high-performance Ethernet/RoCE and InfiniBand fabrics with clear standards for routing, redundancy, and capacity.
  • Automate network provisioning, configuration validation, upgrades, and rollback procedures.
  • Diagnose packet loss, congestion, link failures, and collective communication performance across hosts and switches.
  • Benchmark end-to-end network performance with infrastructure and ML teams, translating workload needs into measurable acceptance criteria.
  • Build monitoring for port health, errors, utilization, congestion, and fabric topology; improve incident response and runbooks.
  • Partner with datacenter operators and hardware vendors on cabling, optics, deployment readiness, and failure resolution.

Skills

Datacenter networking
InfiniBand/RoCE
TCP/IP routing
Linux networking
Python/Ansible automation

Job description

Prime Intellect AI is building an open frontier AI stack and a unified platform for large-scale GPU training, inference, and tooling. You will design and operate networks that connect massive GPU clusters, ensuring reliability and scalable performance across training fabrics, storage networks, and management connectivity.

You will collaborate with hardware and ML teams to set standards, automate provisioning, and monitor fabric health.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Datacenter Networking Engineer for GPU Infra
Staff Datacenter Networking Engineer for GPU Infra

Prime Intellect • United States

On-site
USD 150,000 - 300,000
Staff Datacenter Networking Engineer: Frontier AI GPU
Staff Datacenter Networking Engineer: Frontier AI GPU

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior GPU Cloud Operations Engineer
Senior GPU Cloud Operations Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Senior GPU Data Center Engineer
Senior GPU Data Center Engineer

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Frontier AI GPU Storage Infrastructure Engineer
Frontier AI GPU Storage Infrastructure Engineer

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Networking at Prime Intellect
Member of Technical Staff - Datacenter Networking at Prime Intellect

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000
Staff Network Engineer — GPU Data Center & HPC Networking
Staff Network Engineer — GPU Data Center & HPC Networking

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Networking
Member of Technical Staff - Datacenter Networking

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Networking
Member of Technical Staff - Datacenter Networking

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - Datacenter Networking
Member of Technical Staff - Datacenter Networking

AI Chopping Block • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 300,000