Senior AI Infrastructure Network Engineer (HPC)

Nscale

San Francisco (CA)

On-site

USD 150,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Base + equity
Medical, dental, vision
Flexible PTO

Job summary

Nscale is seeking a Senior Network Engineer – AI Infrastructure to own the health and performance of our Infiniband/RoCE fabrics powering AI workloads. You will lead incident response, drive performance tuning, and collaborate with SREs to build automation. We offer base + equity and a competitive package in a hyperscale environment.

Join a 24/7 operational team focused on reliability, latency, and scalable AI networking, with opportunities to influence fabric design and tooling decisions.

Qualifications

  • 5+ years of experience in network engineering, with at least 3 years operating HPC or large-scale AI interconnect networks.
  • Deep, hands-on operational experience with Infiniband and/or modern RoCE deployments.
  • Proficiency in Linux-based tooling and fabric diagnostics.

Responsibilities

  • Owning the operational health, configuration consistency, and performance tuning of large-scale Infiniband and RoCE fabrics supporting AI and HPC workloads.
  • Leading diagnosis and resolution of complex network incidents (P0/P1) across firmware, kernel drivers, switch hardware, and applications.
  • Driving blameless postmortems and implementing preventative fixes to improve fabric stability and availability.
  • Partnering with SREs to define automation requirements and contribute to network provisioning, validation, and monitoring systems.
  • Collaborating with Network Architecture and Engineering teams to validate fabric designs and enforce routing, congestion control, and firmware baselines.
  • Monitoring fabric utilisation and performance, identifying bottlenecks, and tuning for congestion and latency.
  • Acting as SME for cross-functional teams on high-speed networking and fabric-level performance characteristics.
  • Participating in on-call rotation supporting mission-critical infrastructure.

Skills

Network engineering
RDMA concepts
Linux tooling
Python/Go scripting
High performance networking

Tools

NVIDIA Spectrum switches
ConnectX NICs
Prometheus
Grafana

Job description

Nscale is seeking a Senior Network Engineer – AI Infrastructure to own the health and performance of our Infiniband/RoCE fabrics powering AI workloads. You will lead incident response, drive performance tuning, and collaborate with SREs to build automation. We offer base + equity and a competitive package in a hyperscale environment.

Join a 24/7 operational team focused on reliability, latency, and scalable AI networking, with opportunities to influence fabric design and tooling decisions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Network Engineer
Senior AI Infrastructure Network Engineer

Nscale • New York (NY)

On-site
USD 150,000 - 240,000
Equity
Comprehensive benefits
Retirement plan
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
Senior AI Infra Network Engineer - Infiniband/RoCE
Senior AI Infra Network Engineer - Infiniband/RoCE

Nscale • Seattle (WA)

On-site
USD 150,000 - 240,000
Equity
Medical/dental/vision
Flexible PTO
+2
Principal AI Infrastructure Network Engineer (Equity)
Principal AI Infrastructure Network Engineer (Equity)

Nscale • Seattle (WA), New York (NY), San Francisco (CA), Houston (TX)

On-site
USD 180,000 - 240,000
Base salary + equity
Equity incentives
Dynamic progression plan
Senior Network Engineer — AI/HPC Fabric & Automation
Senior Network Engineer — AI/HPC Fabric & Automation

Nscale • Seattle (WA)

On-site
USD 150,000 - 210,000
Senior AI Infra Networking Engineer | High-Perf GPU Cloud
Senior AI Infra Networking Engineer | High-Perf GPU Cloud

Nscale • United States

Remote
USD 100,000 - 200,000
Principal Network Engineer
Principal Network Engineer

nscaleoperationsukltd • Seattle (WA)

On-site
USD 180,000 - 260,000
Senior Network Engineer, AI Infra & High-Performance Cloud
Senior Network Engineer, AI Infra & High-Performance Cloud

Nscale • San Francisco (CA)

On-site
USD 150,000 - 210,000
Medical insurance
Retirement plan
Flexible PTO
Global HPC Network Engineer for AI Infra
Global HPC Network Engineer for AI Infra

Together • San Francisco (CA)

On-site
USD 190,000 - 280,000
Startup equity
Health insurance
Competitive benefits
Senior Back-End Network Engineer - AI Infrastructure Operations
Senior Back-End Network Engineer - AI Infrastructure Operations

Nscale • Seattle (WA)

On-site
USD 150,000 - 240,000
Equity
Medical/dental/vision
Flexible PTO
+2