GPU Networking Engineer for Large-Scale AI Fabric

Thinking Machines Lab Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 350,000 - 475,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision benefits
Unlimited PTO
Parental leave
Relocation support

Job summary

Thinking Machines Lab Inc. is seeking a network engineer to own the lowest layers of the network stack for large-scale training and inference. You will ensure interconnect reliability across GPU fabrics, debugging NICs, and building instrumentation for faster troubleshooting.

The role is hands-on and cross-stack, working with cloud providers to drive issues to resolution and to keep researchers’ fleets trusted and fast.

Qualifications

  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • Proficiency in at least one backend language (we use Python or Rust).
  • Experience operating large-scale clusters and container orchestration systems (e.g. Kubernetes or Slurm).
  • Comfort operating across the stack and owning projects end-to-end.
  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.
  • A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.

Responsibilities

  • Reason about and validate GPU network fabric design across our deployments.
  • Debug RDMA / RoCEv2 across different NIC vendors. Diagnose collective failures of production NCCL, PFC/ECN tuning, and congestion control behavior.
  • Own NVLink / NVSwitch interconnect - including fabric manager and IMEX health, link and lane errors, and how the GPU fabric interacts with collectives.
  • Build host-level network instrumentation and use Linux tooling to build dashboards and alerts, not just the bug report.
  • Navigate cross-cloud fabric quirks across providers and triage across the NIC, driver, kernel, switch, and workload boundaries.
  • Drive escalations with cloud-provider networking teams, owning issues end-to-end until they're resolved.

Skills

Distributed systems
Backend development
Linux fundamentals
Cross-functional collaboration

Education

Bachelor's degree in CS/Engineering

Tools

Python
Rust
Kubernetes
Slurm
NCCL

Job description

Thinking Machines Lab Inc. is seeking a network engineer to own the lowest layers of the network stack for large-scale training and inference. You will ensure interconnect reliability across GPU fabrics, debugging NICs, and building instrumentation for faster troubleshooting.

The role is hands-on and cross-stack, working with cloud providers to drive issues to resolution and to keep researchers’ fleets trusted and fast.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Network Engineer, Supercomputing
Network Engineer, Supercomputing

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Network Engineer, Supercomputing
Network Engineer, Supercomputing

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision insurance
Unlimited PTO
Paid parental leave
+1
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
GPU Infrastructure Engineer — Scalable AI Training
GPU Infrastructure Engineer — Scalable AI Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1
Senior Network Engineer — AI Infra & HPC Fabric Expert
Senior Network Engineer — AI Infra & HPC Fabric Expert

Nscale • Houston (TX)

On-site
USD 150,000 - 210,000
Competitive benefits package
Flexible paid time off
Parental leave
+1
AI Grid Networking Engineer - GPU Fabric Performance
AI Grid Networking Engineer - GPU Fabric Performance

AMP PBC • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Network Architect for Large-Scale GPU Clusters
Senior AI Network Architect for Large-Scale GPU Clusters

Hamilton Barnes Associates Limited • United States

On-site
USD 220,000 - 350,000
Annual bonus
Equity opportunities
Flexible working arrangements
+1
GPU Networking Architect for AI-Scale Systems
GPU Networking Architect for AI-Scale Systems

NVIDIA • Town of Poland (NY)

On-site
USD 292,000 - 650,000
Comprehensive benefits package
Highly competitive salary
Senior GPU Network Architect for AI Clusters
Senior GPU Network Architect for AI Clusters

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
GPU Network Engineer: RDMA/NVLink at Scale
GPU Network Engineer: RDMA/NVLink at Scale

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision insurance
Unlimited PTO
Paid parental leave
+1