Member of Technical Staff, GPU Systems & Fabric

General Diffusion, Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 250,000

Full time

11 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

General Diffusion, Inc. seeks a Member of Technical Staff to advance multi-GPU systems and fabric performance.

You will design experiments, profile communication paths, and translate results into actionable models for runtime decisions across the GD-X topology. This role owns fabric behavior between devices, collaborates with Compute World Models, Runtime & Placement, and Measurement & Data teams, and maintains regression benchmarks across hardware configurations and workloads to ensure scalable

Qualifications

  • Hands-on experience with multi-GPU or multi-node workloads using collectives (all-reduce, all-gather, reduce-scatter, all-to-all).
  • Deep knowledge of GPU communication stacks (NCCL) and PCIe/NVLink behavior.
  • Ability to design controlled performance experiments and interpret traces and counters.
  • Experience reasoning about topology-aware communication and rank-to-device mapping.
  • Experience debugging hangs, timeouts, and throughput issues in Linux distributed environments.

Responsibilities

  • Design repeatable experiments measuring collective latency, bandwidth, overlap, and tail behavior.
  • Profile end-to-end execution to separate fabric-bound limits from per-device bottlenecks.
  • Develop topology-aware configurations and strategies under representative workloads.
  • Build telemetry and diagnostic artifacts connecting fabric paths to performance outcomes.
  • Maintain regression benchmarks across hardware configurations and software versions.

Skills

Multi-GPU workloads
NCCL / GPU interconnects
Performance experiments
Topology-aware communication
Linux debugging in distributed systems

Tools

NCCL
Profiling tools

Job description

Member of Technical Staff, GPU Systems & Fabric

Improve performance and reliability across multi-GPU systems and their interconnects.

Status Open

Area GPU Systems

Make the communication paths between accelerators measurable, dependable, and useful to the rest of General Diffusion's heterogeneous-compute stack. You will characterize how topology, collectives, data movement, and contention shape multi-GPU execution on GD-X and translate that evidence into actionable interfaces for performance models and runtime decisions. This role owns fabric behavior between devices—not per-device kernel implementation or production placement policy.

01 / The work

What you'll work on
  • Design repeatable experiments that measure collective latency, bandwidth, overlap, and tail behavior across supported multi-GPU and multi-node topologies.
  • Profile end-to-end execution to distinguish communication and topology limits from per-device kernel, memory-capacity, or host-side bottlenecks.
  • Develop and validate topology-aware collective configurations and execution strategies under representative workload shapes, message sizes, and contention.
  • Build telemetry and diagnostic artifacts that connect observed fabric paths, transport choices, failures, and performance outcomes to reproducible runs.
  • Investigate communication stalls, degradation, and configuration mismatches methodically; document bounded mitigations, failure signatures, and reversal criteria.
  • Provide measured fabric capability and cost signals to Compute World Models, Runtime & Placement, and Measurement & Data Infrastructure, with stated assumptions and confidence limits.
  • Maintain regression benchmarks that catch changes in collective performance or reliability across software versions, hardware configurations, and load conditions.

02 / The background

What you bring
  • Hands-on experience improving or operating multi-GPU or multi-node workloads that rely on collectives such as all-reduce, all-gather, reduce-scatter, or all-to-all.
  • Deep practical knowledge of GPU communication stacks - such as NCCL or an equivalent and the behavior of PCIe, NVLink-class links, RDMA-capable NICs, and network fabrics.
  • Ability to construct controlled performance experiments and interpret traces, counters, and timelines rather than attributing slowdowns to a single layer by default.
  • Experience reasoning about topology-aware communication, rank-to-device mapping, rail affinity, synchronization, and the effects of concurrent workloads.
  • Sound systems judgment in Linux-based distributed environments, including debugging hangs, timeouts, throughput regressions, and configuration-dependent failures.

03 / The evidence

What progress looks like
  • Produces a versioned baseline of collective and data-movement behavior across the available testbed topologies, including methodology, variability, and known measurement limits.
  • Delivers a reproducible diagnosis workflow that separates fabric-bound regressions from kernel, memory, host, and runtime causes, demonstrated on representative failure or slowdown cases.
  • Establishes evidence-backed fabric features and cost signals that downstream performance-model and runtime teams can consume, with validation against held-out workload and topology conditions.

04 / In the system

Where this role fits

Owns communication and fabric behavior between devices; Kernels owns per-device primitives and Runtime owns placement.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Engineer, GPU Systems & Fabric
Staff Engineer, GPU Systems & Fabric

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Member of Technical Staff, Distributed Systems & Fleet
Member of Technical Staff, Distributed Systems & Fleet

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 160,000 - 220,000
Member of Technical Staff, GPU & ASIC Performance Modeling
Member of Technical Staff, GPU & ASIC Performance Modeling

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 290,000
Member of Technical Staff, Kernels
Member of Technical Staff, Kernels

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Member of Technical Staff, Measurement & Data Infrastructure
Member of Technical Staff, Measurement & Data Infrastructure

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff, Heterogeneous Runtime & Placement
Member of Technical Staff, Heterogeneous Runtime & Placement

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Applied Researcher – Network Expert
Applied Researcher – Network Expert

Designworks Talent • Bellevue (KY)

Hybrid
USD 120,000 - 190,000
Large-Scale GPU Cluster Engineering Lead (GPU · Cluster · Orchestration)
Large-Scale GPU Cluster Engineering Lead (GPU · Cluster · Orchestration)

NJF Global Holdings Ltd • New York (NY)

On-site
USD 150,000 - 200,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000