GPU Networking Engineer — RDMA/NVLink Fabric

Thinkingmachines

San Francisco (CA)

On-site

USD 350,000 - 475,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Unlimited PTO
Health, dental, and vision benefits
Paid parental leave
Relocation support

Job summary

Thinkingmachines is seeking a network engineer to manage the lowest layers of the network stack pivotal for training and inference. You will ensure interconnect reliability for our large-scale GPU fabrics.

The ideal candidate will have a degree in computer science or engineering, backend programming skills in Python or Rust, and experience with Kubernetes. The position offers a salary range of $350,000 to $475,000 USD.

Benefits include health, dental, vision insurance, and unlimited PTO.

Qualifications

  • Bachelor’s degree or equivalent experience in computer science, engineering, or similar.
  • Experience operating large-scale clusters and container orchestration systems.
  • Comfort operating across the stack and owning projects end-to-end.

Responsibilities

  • Validate GPU network fabric design across deployments.
  • Debug RDMA / RoCEv2 and diagnose production failures.
  • Build network instrumentation and alert systems.

Skills

Backend programming proficiency (Python or Rust)
Experience with container orchestration systems (Kubernetes or Slurm)
Host-level debugging on Linux

Education

Bachelor’s degree in computer science or engineering

Tools

NCCL
CUDA

Job description

Thinkingmachines is seeking a network engineer to manage the lowest layers of the network stack pivotal for training and inference. You will ensure interconnect reliability for our large-scale GPU fabrics.

The ideal candidate will have a degree in computer science or engineering, backend programming skills in Python or Rust, and experience with Kubernetes. The position offers a salary range of $350,000 to $475,000 USD.

Benefits include health, dental, vision insurance, and unlimited PTO.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Network Engineer: RDMA/NVLink at Scale
GPU Network Engineer: RDMA/NVLink at Scale

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision insurance
Unlimited PTO
Paid parental leave
+1
Network Engineer, Supercomputing
Network Engineer, Supercomputing

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision insurance
Unlimited PTO
Paid parental leave
+1
Network Engineer, Supercomputing
Network Engineer, Supercomputing

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Unlimited PTO
Health, dental, and vision benefits
Paid parental leave
+1
GPU Networking Engineer — RDMA & Distributed Inference
GPU Networking Engineer — RDMA & Distributed Inference

Baseten • United States

Remote
USD 180,000 - 240,000
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Senior GPU AI Network Architect
Senior GPU AI Network Architect

Sesterce Group • San Francisco (CA)

On-site
USD 120,000 - 160,000
GPU Networking Engineer - RDMA & Distributed Inference
GPU Networking Engineer - RDMA & Distributed Inference

Baseten • San Francisco (CA)

On-site
USD 180,000 - 260,000
Competitive compensation
100% medical, dental, and vision insurance
Generous PTO policy
+3
Senior Software Engineer, Fabric Networking - GPU
Senior Software Engineer, Fabric Networking - GPU

NVIDIA • Arizona

On-site
USD 152,000 - 288,000
Equity
Generous benefits package
Principal Architect, AI Networking
Principal Architect, AI Networking

NVIDIA • Town of Texas (WI)

On-site
USD 272,000 - 432,000
Principal Architect, AI Networking
Principal Architect, AI Networking

NVIDIA • Oregon (WI)

On-site
USD 272,000 - 432,000