Staff Datacenter Networking Engineer: Frontier AI GPU

Prime Intellect

San Francisco (CA)

On-site

USD 150,000 - 300,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Prime Intellect is building the open superintelligence stack and operates a datacenter network backbone for large GPU clusters. The role focuses on designing, deploying, and maintaining scalable network fabrics to support training and inference workloads at frontier scale.

You will work on Ethernet and InfiniBand/RoCE networks, automate provisioning and monitoring, and partner with hardware vendors to ensure reliability and performance of distributed workloads.

Qualifications

  • 3+ years of production datacenter networking experience.
  • Strong understanding of Ethernet, TCP/IP, routing, switching and redundant network design.

Responsibilities

  • Design and deploy scalable datacenter network topologies for GPU training, inference, storage, and management traffic.
  • Configure and operate high-performance Ethernet/RoCE and InfiniBand fabrics with clear standards for routing, redundancy, and capacity.
  • Automate network provisioning, configuration validation, upgrades, and rollback procedures.
  • Diagnose packet loss, congestion, link failures, and performance across hosts and switches.
  • Benchmark end-to-end network performance with infra and ML teams; translate workload needs into acceptance criteria.
  • Build monitoring for port health, errors, utilization, congestion, and fabric topology; improve incident response.
  • Partner with datacenter operators and hardware vendors on cabling, optics, deployment readiness, and failure resolution.

Skills

Datacenter networking
Networking fundamentals
InfiniBand/RoCE
Automation (Python/Ansible)

Tools

InfiniBand
RoCE
NIC drivers

Job description

Prime Intellect is building the open superintelligence stack and operates a datacenter network backbone for large GPU clusters. The role focuses on designing, deploying, and maintaining scalable network fabrics to support training and inference workloads at frontier scale.

You will work on Ethernet and InfiniBand/RoCE networks, automate provisioning and monitoring, and partner with hardware vendors to ensure reliability and performance of distributed workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Datacenter Networking
Member of Technical Staff - Datacenter Networking

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Storage Systems Engineer for Frontier AI
Storage Systems Engineer for Frontier AI

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
GPU Cloud Infrastructure Engineer
GPU Cloud Infrastructure Engineer

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior Cluster Infra Architect for Frontier AI — Equity
Senior Cluster Infra Architect for Frontier AI — Equity

RadixArk • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Member of Technical Staff - Datacenter Operations
Member of Technical Staff - Datacenter Operations

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Senior AI Infra Network Engineer (Datacenter/GPU) (Equity)
Senior AI Infra Network Engineer (Datacenter/GPU) (Equity)

QumulusAI • United States

On-site
USD 100,000 - 130,000
GPU Networking Engineer for Large-Scale AI Fabric
GPU Networking Engineer for Large-Scale AI Fabric

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1