GPU Network Engineer: RDMA/NVLink at Scale

Thinking Machines Lab

San Francisco (CA)

On-site

USD 350,000 - 475,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health, dental, and vision insurance
Unlimited PTO
Paid parental leave
Relocation support

Job summary

Thinking Machines Lab is seeking a Network Engineer in San Francisco to manage and improve our GPU network fabric. The role requires in-depth knowledge of large-scale deployments and the ability to debug complex network issues. A collaborative environment is emphasized, where initiative and effective communication with cloud providers are key.

The position offers a competitive salary ranging from $350,000 to $475,000 per year, depending on skills and experience. Benefits include unlimited PTO and health coverage, alongside visa sponsorship.

Qualifications

  • Degree or equivalent in computer science, engineering, or a related field.
  • Experience with RDMA/RoCE and debugging networking issues.
  • Ability to work across multiple cloud providers.

Responsibilities

  • Manage the reliability of GPU network fabric across deployments.
  • Debug network issues and automate tooling for efficient debugging.
  • Collaborate with cloud providers to resolve connectivity issues.

Skills

Backend language proficiency (Python or Rust)
Experience with large-scale clusters
Comfortable across the tech stack
Collaboration skills
Initiative and action bias

Education

Bachelor’s degree in computer science or related field

Tools

Linux debugging tools
Kubernetes or Slurm
CUDA/NCCL

Job description

Thinking Machines Lab is seeking a Network Engineer in San Francisco to manage and improve our GPU network fabric. The role requires in-depth knowledge of large-scale deployments and the ability to debug complex network issues. A collaborative environment is emphasized, where initiative and effective communication with cloud providers are key.

The position offers a competitive salary ranging from $350,000 to $475,000 per year, depending on skills and experience. Benefits include unlimited PTO and health coverage, alongside visa sponsorship.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Networking Engineer — RDMA/NVLink Fabric
GPU Networking Engineer — RDMA/NVLink Fabric

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Unlimited PTO
Health, dental, and vision benefits
Paid parental leave
+1
Network Engineer, Supercomputing
Network Engineer, Supercomputing

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision insurance
Unlimited PTO
Paid parental leave
+1
Network Engineer, Supercomputing
Network Engineer, Supercomputing

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Unlimited PTO
Health, dental, and vision benefits
Paid parental leave
+1
Senior Network Engineer - DGX Cloud
Senior Network Engineer - DGX Cloud

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 168,000 - 333,000
Equity
Benefits
GPU Systems Engineer
GPU Systems Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
GPU Systems Engineer
GPU Systems Engineer

Iceberg • New York (NY)

On-site
USD 200,000 - 300,000
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Senior Software Engineer, Fabric Networking - GPU
Senior Software Engineer, Fabric Networking - GPU

NVIDIA • Arizona

On-site
USD 152,000 - 288,000
Equity
Generous benefits package
Senior Data Center Network Engineer – GPU Clusters
Senior Data Center Network Engineer – GPU Clusters

Baseten • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation, including meaningful equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+4
Senior Network Engineer - DGX Cloud
Senior Network Engineer - DGX Cloud

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 168,000 - 334,000