Member of Technical Staff, GPU & ASIC Performance Modeling

General Diffusion, Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 290,000

Full time

10 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

General Diffusion, Inc. in San Francisco seeks a Member of Technical Staff to develop portable measurements and predictive performance models for external GPUs and ASICs, covering compute, memory, data movement, and contention.

You will produce traceable, evidence-backed capability profiles for compute-world-model, compiler, and runtime partners. The role emphasizes calibrating models to measurements, designing repeatable experiments, and collaborating with GPU Systems & Fabric on multi-device

Qualifications

  • Experience in accelerator architecture or performance modeling grounded in measured GPU, inference-ASIC, or comparable heterogeneous-system behavior.
  • Ability to design benchmarks that combine focused microkernels with representative workloads, control confounding variables, and make results reproducible.
  • Demonstrated model-to-measurement correlation work, including calibration, error analysis, and clear treatment of uncertainty under workload or architecture shift.
  • Fluency with performance-profiling methods and the interaction of arithmetic intensity, memory hierarchy, occupancy or utilization, and data movement.
  • Vendor-neutral hardware–software judgment: able to compare capabilities through evidence rather than feature lists, and communicate limits to systems and ML collaborators.

Responsibilities

  • Design controlled microbenchmarks and workload-level experiments that isolate compute throughput, memory-hierarchy behavior, data movement, and contention on external GPU and ASIC targets.
  • Build repeatable characterization protocols that record workload, hardware, software, configuration, and measurement conditions so comparisons across targets are traceable.
  • Develop and validate performance models against measured behavior; quantify prediction error, uncertainty, and the conditions under which a model transfers or ceases to transfer.
  • Use profiler evidence and end-to-end measurements to distinguish compute, memory, communication, and interference limits rather than treating peak specifications as attainable performance.
  • Publish decision-ready capability profiles and methodological caveats for compute-world-model, compiler, and runtime partners, including the evidence behind each conclusion.
  • Partner with GPU Systems & Fabric on multi-device measurements and incorporate relevant data-movement characteristics into profiles, while leaving collective and fabric optimization to that team.

Skills

Accelerator architecture
Performance modeling
Benchmark design
Profiling methods
Data movement

Job description

Member of Technical Staff, GPU & ASIC Performance Modeling

Understand external GPUs and ASICs through portable measurements and predictive performance models.

Status Open

Area Hardware Research

Develop portable measurements and predictive performance models that explain how externally available GPUs and ASICs behave under representative workloads, including compute, memory, data movement, and contention. Turn that evidence into calibrated capability profiles that help compute-world-model, compiler, and runtime teams reason about unlike targets—without owning kernel implementation, compiler lowering, placement policy, or silicon design.

01 / The work

What you’ll work on
  • Design controlled microbenchmarks and workload-level experiments that isolate compute throughput, memory-hierarchy behavior, data movement, and contention on external GPU and ASIC targets.
  • Build repeatable characterization protocols that record workload, hardware, software, configuration, and measurement conditions so comparisons across targets are traceable.
  • Develop and validate performance models against measured behavior; quantify prediction error, uncertainty, and the conditions under which a model transfers or ceases to transfer.
  • Use profiler evidence and end-to-end measurements to distinguish compute, memory, communication, and interference limits rather than treating peak specifications as attainable performance.
  • Publish decision-ready capability profiles and methodological caveats for compute-world-model, compiler, and runtime partners, including the evidence behind each conclusion.
  • Partner with GPU Systems & Fabric on multi-device measurements and incorporate relevant data-movement characteristics into profiles, while leaving collective and fabric optimization to that team.
02 / The background
What you bring
  • Experience in accelerator architecture or performance modeling grounded in measured GPU, inference-ASIC, or comparable heterogeneous-system behavior.
  • Ability to design benchmarks that combine focused microkernels with representative workloads, control confounding variables, and make results reproducible.
  • Demonstrated model-to-measurement correlation work, including calibration, error analysis, and clear treatment of uncertainty under workload or architecture shift.
  • Fluency with performance-profiling methods and the interaction of arithmetic intensity, memory hierarchy, occupancy or utilization, and data movement.
  • Vendor-neutral hardware–software judgment: able to compare capabilities through evidence rather than feature lists, and communicate limits to systems and ML collaborators.
03 / The evidence
What progress looks like
  • A traceable characterization suite produces repeatable capability profiles across selected external targets, with workload and environment provenance sufficient for another engineer to reproduce the comparison.
  • Validation reports compare model estimates with held-out measurements, quantify error and uncertainty by workload regime, and document transfer limits rather than obscuring mismatches.
  • Compiler, runtime, and research partners receive evidence-backed profiles that identify the relevant bottleneck or compatibility caveat for a target and cite the underlying measurements.
04 / In the system
Where this role fits

This is characterization of external silicon, not RTL chip design or ownership of a vendor roadmap.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior External GPU & ASIC Performance Modeling Engineer
Senior External GPU & ASIC Performance Modeling Engineer

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 290,000
Principal Competitive CPU Microarchitecture & Platform Performance Engineer
Principal Competitive CPU Microarchitecture & Platform Performance Engineer

AMD • Austin (TX)

On-site
USD 190,000 - 260,000
Member of Technical Staff, GPU Systems & Fabric
Member of Technical Staff, GPU Systems & Fabric

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
ASIC Engineer
ASIC Engineer

Anysilicon • Austin (TX), Northern (KY)

On-site
USD 150,000 - 210,000
Silicon Performance Architect, Reality Labs
Silicon Performance Architect, Reality Labs

Meta • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Kernels
Member of Technical Staff, Kernels

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
ASIC Engineer, Performance Architecture and Modeling
ASIC Engineer, Performance Architecture and Modeling

Meta • Sunnyvale (CA), Austin (TX)

On-site
USD 180,000 - 240,000
Pre-Silicon CPU Architecture & Platform Performance Engineer
Pre-Silicon CPU Architecture & Platform Performance Engineer

Advanced Micro Devices • Town of Texas (WI)

On-site
USD 140,000 - 190,000
Performance Modeling Lead
Performance Modeling Lead

OpenAI • Los Angeles (CA)

On-site
USD 130,000 - 180,000
Relocation assistance
Hybrid work model
System Architect
System Architect

AheadComputing, Inc. • Northern (KY)

Hybrid
USD 140,000 - 210,000
Competitive compensation
Flexible and inclusive culture
Professional growth opportunities
+1