Senior AI Infra Network Engineer - Infiniband/RoCE

Nscale

Seattle (WA)

On-site

USD 150,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Medical/dental/vision
Flexible PTO
Parental leave
Retirement plan

Job summary

Nscale, a GPU cloud provider for AI, seeks a Senior Network Engineer – AI Infrastructure in Seattle to own health and performance of Infiniband and RoCE fabrics for AI/HPC workloads.

You will diagnose complex incidents, drive improvements, and collaborate with SREs to automate provisioning and monitoring. A strong Linux, RDMA, and TCP/IP foundation is essential for this role, with on-call responsibilities included.

Qualifications

  • 5+ years of experience in network engineering, with at least 3 years operating HPC or large-scale AI interconnect networks.
  • Deep, hands-on operational experience with Infiniband and/or modern RoCE deployments.
  • Expert understanding of RDMA concepts, protocols, and troubleshooting techniques.
  • Strong fundamentals in data centre networking, including TCP/IP, BGP, OSPF, and leaf-spine architectures.
  • Proven ability to troubleshoot complex network issues using Linux-based tooling and fabric diagnostics.
  • Proficiency in Python, Go, or shell scripting for automation, data analysis, or configuration management.
  • Experience working in a 24/7 operational environment with a strong focus on reliability and toil reduction.

Responsibilities

  • Owning the operational health, configuration consistency, and performance tuning of large-scale Infiniband and RoCE fabrics supporting AI and HPC workloads.
  • Leading the diagnosis and resolution of complex network incidents (P0/P1), spanning firmware, kernel drivers, switch hardware, and application or middleware layers.
  • Driving blameless postmortems and implementing preventative fixes to improve long-term fabric stability and availability.
  • Partnering with SREs to define requirements for automation and tooling, and contributing where appropriate to network provisioning, validation, and monitoring systems.
  • Collaborating with Network Architecture and Engineering teams to validate fabric designs and enforce standards for routing, congestion control, and firmware baselines.
  • Monitoring fabric utilisation and performance, identifying bottlenecks, and tuning for congestion, microbursts, and predictable latency.
  • Acting as a subject matter expert for cross-functional teams on high-speed networking, RDMA behaviour, and fabric-level performance characteristics.
  • Participating in an on-call rotation supporting mission-critical, customer-facing infrastructure

Skills

Infiniband
RoCE
RDMA
Linux
Python
Go
Shell
Networking
BGP/OSPF
TCP/IP

Tools

Prometheus
Grafana
NCCL

Job description

Nscale, a GPU cloud provider for AI, seeks a Senior Network Engineer – AI Infrastructure in Seattle to own health and performance of Infiniband and RoCE fabrics for AI/HPC workloads.

You will diagnose complex incidents, drive improvements, and collaborate with SREs to automate provisioning and monitoring. A strong Linux, RDMA, and TCP/IP foundation is essential for this role, with on-call responsibilities included.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Network Engineer
Senior AI Infrastructure Network Engineer

Nscale • New York (NY)

On-site
USD 150,000 - 240,000
Equity
Comprehensive benefits
Retirement plan
Senior AI Infrastructure Network Engineer (HPC)
Senior AI Infrastructure Network Engineer (HPC)

Nscale • San Francisco (CA)

On-site
USD 150,000 - 240,000
Base + equity
Medical, dental, vision
Flexible PTO
Principal AI Infra Network Engineer (Infiniband/RDMA)
Principal AI Infra Network Engineer (Infiniband/RDMA)

Socket.dev • Houston (TX)

On-site
USD 180,000 - 270,000
Equity
Competitive base package
Career growth opportunities
Principal AI Infrastructure Network Engineer (Equity)
Principal AI Infrastructure Network Engineer (Equity)

Nscale • Seattle (WA), New York (NY), San Francisco (CA), Houston (TX)

On-site
USD 180,000 - 240,000
Base salary + equity
Equity incentives
Dynamic progression plan
Senior AI Infra Networking Engineer | High-Perf GPU Cloud
Senior AI Infra Networking Engineer | High-Perf GPU Cloud

Nscale • United States

Remote
USD 100,000 - 200,000
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Senior Network Engineer — AI/HPC Fabric & Automation
Senior Network Engineer — AI/HPC Fabric & Automation

Nscale • Seattle (WA)

On-site
USD 150,000 - 210,000
Senior InfiniBand Network Engineer for AI Scale
Senior InfiniBand Network Engineer for AI Scale

Neura Market • New York (NY)

Hybrid
USD 170,000 - 210,000
Health Coverage
Equity/RSUs
401(k) matching
+8
Senior Network Engineer, AI Infra & High-Performance Cloud
Senior Network Engineer, AI Infra & High-Performance Cloud

Nscale • San Francisco (CA)

On-site
USD 150,000 - 210,000
Medical insurance
Retirement plan
Flexible PTO
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)

Sciforium • San Francisco (CA)

On-site
USD 170,000 - 230,000