GPU Networking Engineer — RDMA & Distributed Inference

Baseten

United States

Remote

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Baseten is building the global operating system for distributed AI workloads. We are seeking foundational engineers to lead our GPU networking efforts, integrating RDMA into the inference stack and enabling scalable, low-latency distributed serving of large models.

You will architect software fabric unifying thousands of GPUs, tapping into Open-Ecosystem while building from scratch when needed to unlock next-gen performance and startup speeds.

Responsibilities

  • Make RDMA First-Class: You will work on integrating RDMA/RoCE/InfiniBand capabilities directly into our inference stack, helping us move beyond TCP/IP to unlock order-of-magnitude improvements in bandwidth and latency.
  • Optimize Distributed Inference: You will implement and tune the networking layers necessary for efficient Disaggregated KV Cache Offload and WideEP, ensuring seamless communication across NVLink and InfiniBand for our MoE models.
  • Enable Serverless-Grade Startup Speeds for LLMs: You will work deeply with checkpointing and storage mechanisms to enable sub-10-second startup for trillion-parameter models.
  • Deep-Dive into Hardware: You will characterize and validate networking performance on bleeding-edge clusters (H100/H200, B200/B300, GB200/300 NVL72), writing the acceptance tests that ensure our hardware delivers peak achievable throughput and minimal latency.
  • Build Observability: You will design the tools that let us visualize packet flow, congestion, and effective bandwidth across the GPU interconnects, helping us diagnose complex distributed system behaviors.

Skills

RDMA
InfiniBand
NVLink
Distributed systems
Networking stack
C/C++
Performance optimization

Job description

Baseten is building the global operating system for distributed AI workloads. We are seeking foundational engineers to lead our GPU networking efforts, integrating RDMA into the inference stack and enabling scalable, low-latency distributed serving of large models.

You will architect software fabric unifying thousands of GPUs, tapping into Open-Ecosystem while building from scratch when needed to unlock next-gen performance and startup speeds.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Networking Engineer for Distributed Inference & RDMA
GPU Networking Engineer for Distributed Inference & RDMA

The Consensus • New York (NY)

On-site
USD 120,000 - 160,000
Competitive compensation, including equity
100% medical, dental, and vision insurance
Flexible PTO policy
+3
RDMA-First GPU Networking & Distributed Inference Engineer
RDMA-First GPU Networking & Distributed Inference Engineer

Baseten • New York (NY)

On-site
USD 185,000 - 250,000
GPU Networking Engineer - RDMA & Distributed Inference
GPU Networking Engineer - RDMA & Distributed Inference

Baseten • San Francisco (CA)

On-site
USD 180,000 - 260,000
Competitive compensation
100% medical, dental, and vision insurance
Generous PTO policy
+3
Software Engineer - GPU Networking & Distributed Systems
Software Engineer - GPU Networking & Distributed Systems

Baseten • New York (NY)

On-site
USD 185,000 - 250,000
Competitive compensation
100% medical coverage
Generous PTO policy
+3
Senior GPU AI Network Architect
Senior GPU AI Network Architect

Sesterce Group • San Francisco (CA)

On-site
USD 120,000 - 160,000
GPU Networking Engineer — RDMA/NVLink Fabric
GPU Networking Engineer — RDMA/NVLink Fabric

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Unlimited PTO
Health, dental, and vision benefits
Paid parental leave
+1
Software Engineer — GPU Networking & Distributed Systems
Software Engineer — GPU Networking & Distributed Systems

The Consensus • New York (NY)

On-site
USD 120,000 - 160,000
Competitive compensation, including equity
100% medical, dental, and vision insurance
Flexible PTO policy
+3
Software Engineer GPU Networking & Distributed Systems
Software Engineer GPU Networking & Distributed Systems

Baseten • San Francisco (CA)

On-site
USD 180,000 - 260,000
Competitive compensation
100% medical, dental, and vision insurance
Generous PTO policy
+3
Principal Engineer - AI Networking
Principal Engineer - AI Networking

Ll Oefentherapie • Seattle (WA)

On-site
USD 100,000 - 130,000
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)
Senior GPU Cluster Networking Engineer (RDMA/InfiniBand)

Sciforium • San Francisco (CA)

On-site
USD 170,000 - 230,000