Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)

Sciforium

San Francisco (CA)

On-site

USD 170,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Sciforium is building an AI infrastructure platform and needs a Senior Network Engineer to own the networking stack for GPU clusters. You will design, bring up, and operate the complete fabric from InfiniBand/RoCE to data center perimeter and cloud connectivity.

You will lead fabric validation, optimize performance, and manage security, routing, and network automation across multi-vendor environments. This role drives science-grade connectivity for scale.

Qualifications

  • 7+ years designing and operating production data center networks, including at least one large-scale HPC/AI cluster fabric (InfiniBand or RoCE v2).
  • Expert-level routing and switching: BGP, OSPF, ECMP, VRF across multiple vendor platforms.
  • Deep RDMA expertise: RoCE v2 tuning and InfiniBand fabric management.
  • Production network security experience: firewalls, VPNs, segmentation.
  • Network automation proficiency: Python plus Ansible/Nornir/NAPALM with Git workflows.

Responsibilities

  • Fabric Architecture & Cluster Network Design for GPU clusters, including topology and oversubscription analysis.
  • RDMA & Performance Engineering: configure RoCE v2 and InfiniBand fabrics, validate with performance tests.
  • Production Network Operations & Security: manage routing, perimeter security, NetBox-based config generation, observability.
  • Cross-Cluster & Cloud Connectivity: design inter-site links and hybrid cloud connectivity.

Skills

Network design
Routing & switching
RDMA expertise
Security
Network automation
Cloud networking
GPUDirect / NCCL mapping

Tools

Ansible
Nornir
NAPALM
Git / version control
Prometheus/Grafana

Job description

Sciforium is building an AI infrastructure platform and needs a Senior Network Engineer to own the networking stack for GPU clusters. You will design, bring up, and operate the complete fabric from InfiniBand/RoCE to data center perimeter and cloud connectivity.

You will lead fabric validation, optimize performance, and manage security, routing, and network automation across multi-vendor environments. This role drives science-grade connectivity for scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Cluster Engineer, Networking
GPU Cluster Engineer, Networking

Sciforium • San Francisco (CA)

On-site
USD 170,000 - 230,000
Senior GPU Cluster Engineer for AI Infrastructure
Senior GPU Cluster Engineer for AI Infrastructure

Sciforium • San Francisco (CA)

On-site
USD 150,000 - 220,000
Medical, dental, and vision insurance
401k plan
Daily lunch, snacks, and beverages
+2
Senior AI Infra Network Engineer - Infiniband/RoCE
Senior AI Infra Network Engineer - Infiniband/RoCE

Nscale • Seattle (WA)

On-site
USD 150,000 - 240,000
Equity
Medical/dental/vision
Flexible PTO
+2
GPU Cloud Network SRE - InfiniBand & RoCE Expert
GPU Cloud Network SRE - InfiniBand & RoCE Expert

Doist • Austin (TX), San Jose (CA)

Hybrid
USD 140,000 - 210,000
GPU Cloud Network SRE: InfiniBand & RoCE Expert
GPU Cloud Network SRE: InfiniBand & RoCE Expert

Triwill Group • United States

Remote
USD 150,000 - 230,000
GPU DC East-West Network SRE: InfiniBand & RoCE Expert
GPU DC East-West Network SRE: InfiniBand & RoCE Expert

Bitdeer • San Jose (CA)

On-site
USD 140,000 - 210,000
Senior Network Support Engineer - InfiniBand & HPC
Senior Network Support Engineer - InfiniBand & HPC

NVIDIA • Germany (OH)

On-site
USD 120,000 - 180,000
Senior InfiniBand Network Engineer for AI Scale
Senior InfiniBand Network Engineer for AI Scale

Neura Market • New York (NY)

Hybrid
USD 170,000 - 210,000
Health Coverage
Equity/RSUs
401(k) matching
+8
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Senior AI Infrastructure Network Engineer
Senior AI Infrastructure Network Engineer

Nscale • New York (NY)

On-site
USD 150,000 - 240,000
Equity
Comprehensive benefits
Retirement plan