Staff Engineer, GPU Systems & Fabric

General Diffusion, Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 250,000

Full time

10 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

General Diffusion, Inc. seeks a Member of Technical Staff to advance multi-GPU systems and fabric performance.

You will design experiments, profile communication paths, and translate results into actionable models for runtime decisions across the GD-X topology. This role owns fabric behavior between devices, collaborates with Compute World Models, Runtime & Placement, and Measurement & Data teams, and maintains regression benchmarks across hardware configurations and workloads to ensure scalable

Qualifications

  • Hands-on experience with multi-GPU or multi-node workloads using collectives (all-reduce, all-gather, reduce-scatter, all-to-all).
  • Deep knowledge of GPU communication stacks (NCCL) and PCIe/NVLink behavior.
  • Ability to design controlled performance experiments and interpret traces and counters.
  • Experience reasoning about topology-aware communication and rank-to-device mapping.
  • Experience debugging hangs, timeouts, and throughput issues in Linux distributed environments.

Responsibilities

  • Design repeatable experiments measuring collective latency, bandwidth, overlap, and tail behavior.
  • Profile end-to-end execution to separate fabric-bound limits from per-device bottlenecks.
  • Develop topology-aware configurations and strategies under representative workloads.
  • Build telemetry and diagnostic artifacts connecting fabric paths to performance outcomes.
  • Maintain regression benchmarks across hardware configurations and software versions.

Skills

Multi-GPU workloads
NCCL / GPU interconnects
Performance experiments
Topology-aware communication
Linux debugging in distributed systems

Tools

NCCL
Profiling tools

Job description

General Diffusion, Inc. seeks a Member of Technical Staff to advance multi-GPU systems and fabric performance.

You will design experiments, profile communication paths, and translate results into actionable models for runtime decisions across the GD-X topology. This role owns fabric behavior between devices, collaborates with Compute World Models, Runtime & Placement, and Measurement & Data teams, and maintains regression benchmarks across hardware configurations and workloads to ensure scalable

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior External GPU & ASIC Performance Modeling Engineer
Senior External GPU & ASIC Performance Modeling Engineer

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 290,000
Member of Technical Staff, GPU Systems & Fabric
Member of Technical Staff, GPU Systems & Fabric

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Staff Engineer, Distributed Systems & Fleet Reliability
Staff Engineer, Distributed Systems & Fleet Reliability

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 160,000 - 220,000
Staff Engineer: Heterogeneous Runtime & Placement
Staff Engineer: Heterogeneous Runtime & Placement

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Kernel Performance Engineer (CUDA/Triton)
Kernel Performance Engineer (CUDA/Triton)

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Staff C++ Engineer
Staff C++ Engineer

Glocomms • San Francisco (CA)

On-site
USD 190,000 - 270,000
Staff Engineer, Distributed GPU Clusters
Staff Engineer, Distributed GPU Clusters

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Staff C++ Engineer: High-Performance GPU Systems
Staff C++ Engineer: High-Performance GPU Systems

Glocomms • San Francisco (CA)

On-site
USD 190,000 - 270,000
Staff Network Engineer — GPU Data Center & HPC Networking
Staff Network Engineer — GPU Data Center & HPC Networking

Matcha • Northern (KY)

Hybrid
USD 150,000 - 300,000
Senior GPU Fabric Architect for AI Cloud Infra
Senior GPU Fabric Architect for AI Cloud Infra

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 180,000 - 240,000