AI Grid Networking Engineer - GPU Fabric Performance

AMP PBC

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AMP PBC is building the largest independent AI fleet. You will own the east-west network across clusters, bringing up the fabric from switch access to production-grade performance for high-scale GPU training workloads.

This founding role requires hands-on deployment, tuning, and leadership to ensure robust, scalable network infrastructure. You will work with RoCEv2/RDMA at scale, diverse switch platforms, and NCCL benchmarking, shaping the fabric for frontier AI labs.

Qualifications

  • Experience bringing up and tuning large GPU training fabrics.
  • Direct hands-on RoCEv2/RDMA in production for 1,000+ GPUs.
  • Deep knowledge of switch platforms (NVIDIA Spectrum-X, Arista, Cisco, Juniper).
  • Comfort with NCCL and benchmarking.

Responsibilities

  • Own east-west fabric bringup and tuning across clusters, starting with a B300 deployment and extending to the full Grid
  • Take clusters from racked and cabled to production performance: switch access, fabric configuration, validation, and handoff to customers running real training workloads
  • Tune lossless Ethernet end to end, including PFC, ECN and DCQCN congestion control, buffer allocation, QoS classes, hashing and traffic isolation
  • Diagnose and fix the failure modes that quietly destroy training throughput: link health, packet loss, congestion collapse, low-entropy and bursty AI traffic patterns
  • Benchmark and validate fabric performance against real collective operations, and defend the numbers our customers depend on
  • Work across heterogeneous hardware by design. Different providers, different sites, different switch vendors and different silicon, with NVIDIA reference architectures as a floor rather than an answer
  • Build the tooling and runbooks that make cluster onboarding fast and repeatable, so the fleet can scale without the process scaling with it
  • Partner directly with frontier AI labs on the performance of the compute they are training on
  • Hire and lead the network team as the Grid grows

Skills

RoCEv2 RDMA
Switch platforms depth
NVIDIA Spectrum-X
Leaf-spine topology
NCCL benchmarking
Team leadership

Education

Bachelor's degree

Tools

Arista EOS
Cisco Nexus
Juniper
SONiC
Cumulus Linux

Job description

AMP PBC is building the largest independent AI fleet. You will own the east-west network across clusters, bringing up the fabric from switch access to production-grade performance for high-scale GPU training workloads.

This founding role requires hands-on deployment, tuning, and leadership to ensure robust, scalable network infrastructure. You will work with RoCEv2/RDMA at scale, diverse switch platforms, and NCCL benchmarking, shaping the fabric for frontier AI labs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Network Performance and Reliability Engineer
AI Network Performance and Reliability Engineer

AMP PBC • San Francisco (CA)

On-site
USD 180,000 - 240,000
Principal AI Networking Engineer for GPU Clusters
Principal AI Networking Engineer for GPU Clusters

The Consensus • United States

On-site
USD 248,000 - 269,000
GPU Networking Engineer for Large-Scale AI Fabric
GPU Networking Engineer for Large-Scale AI Fabric

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
AI Network Architect: Scalable GPU Fabrics & Automation
AI Network Architect: Scalable GPU Fabrics & Automation

Supermicro • Wayne (CA)

On-site
USD 200,000 - 220,000
AI Network Architect: GPU Fabric & High-Speed Switches
AI Network Architect: GPU Fabric & High-Speed Switches

Super Micro Computer Spain, S.L. • San Jose (CA)

On-site
USD 200,000 - 220,000
Bonus programs
Equity award programs
Senior AI Infra Network Engineer (Datacenter/GPU) (Equity)
Senior AI Infra Network Engineer (Datacenter/GPU) (Equity)

QumulusAI • United States

On-site
USD 100,000 - 130,000
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Lead AI Networking & Infrastructure Engineer
Lead AI Networking & Infrastructure Engineer

Upscale AI • United States

On-site
USD 248,000 - 269,000
Remote Data Center Network Engineer — AI & GPU-Scale Fabric
Remote Data Center Network Engineer — AI & GPU-Scale Fabric

Matrix USA • United States

On-site
USD 120,000 - 180,000