GPU Networking Engineer for Distributed Inference & RDMA

The Consensus

New York (NY)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation, including equity
100% medical, dental, and vision insurance
Flexible PTO policy
Paid parental leave
Fertility and family-building stipend
Company-facilitated 401(k)

Job summary

The Consensus is seeking experienced foundational engineers to optimize networking in AI infrastructures. You will play a crucial role in integrating and improving high-performance networking technologies to enhance distributed AI capabilities.

The role demands deep technical expertise in C++ or Python, a grasp on high-performance networking protocols, and the ability to optimize models for unprecedented performance metrics. Join a pioneering team in shaping the future of AI technology.

Qualifications

  • Deep experience with high-performance networking protocols (InfiniBand, RoCE v2) required.
  • Fluency in C++ or Python with a deep understanding of modern NVIDIA architectures.
  • Ability to debug NVLink topology issues and customize solutions.

Responsibilities

  • Integrate RDMA/RoCE/InfiniBand capabilities into the inference stack.
  • Optimize distributed inference for efficient communication across models.
  • Design tools for visualizing packet flow and effective bandwidth.

Skills

High-performance networking protocols
C++ or Python
Memory hierarchy optimization
Custom communication kernels

Tools

NCCL
NVSHMEM
TensorRT-LLM
GPUDirect Storage

Job description

The Consensus is seeking experienced foundational engineers to optimize networking in AI infrastructures. You will play a crucial role in integrating and improving high-performance networking technologies to enhance distributed AI capabilities.

The role demands deep technical expertise in C++ or Python, a grasp on high-performance networking protocols, and the ability to optimize models for unprecedented performance metrics. Join a pioneering team in shaping the future of AI technology.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Networking Engineer - RDMA & Distributed Inference
GPU Networking Engineer - RDMA & Distributed Inference

Baseten • San Francisco (CA)

On-site
USD 180,000 - 260,000
Competitive compensation
100% medical, dental, and vision insurance
Generous PTO policy
+3
GPU Networking Engineer — RDMA & Distributed Inference
GPU Networking Engineer — RDMA & Distributed Inference

Baseten • United States

Remote
USD 180,000 - 240,000
RDMA-First GPU Networking & Distributed Inference Engineer
RDMA-First GPU Networking & Distributed Inference Engineer

Baseten • New York (NY)

On-site
USD 185,000 - 250,000
Principal Engineer - AI Networking
Principal Engineer - AI Networking

Ll Oefentherapie • Seattle (WA)

On-site
USD 100,000 - 130,000
Senior GPU AI Network Architect
Senior GPU AI Network Architect

Sesterce Group • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior AI Networking & Performance Engineer
Senior AI Networking & Performance Engineer

NVIDIA • Town of Texas (WI)

On-site
USD 272,000 - 432,000
Senior AI Networking Engineer | Embedded RDMA/CUDA Expert
Senior AI Networking Engineer | Embedded RDMA/CUDA Expert

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Senior AI Networking & Performance Engineer
Senior AI Networking & Performance Engineer

NVIDIA • Colorado

On-site
USD 272,000 - 432,000
GPU Network Engineer
GPU Network Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD <240,000
Senior AI Networking & Performance Architect (GPU/DL)
Senior AI Networking & Performance Architect (GPU/DL)

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Competitive salary
Generous benefits package
Equity eligibility