Member of Technical Staff, Kernels

General Diffusion, Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 280,000

Full time

10 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

General Diffusion, Inc. in San Francisco seeks a senior engineer to own kernel hot paths across heterogeneous accelerators, turning profiler evidence into portable, numerically correct kernels.

You will profile workloads, design and tune kernels in CUDA/Triton, optimize data layouts and launch configurations, and maintain regression tests and reference implementations to prove performance and correctness. The role emphasizes portability, clear documentation for compiler and runtime partners, and

Qualifications

  • Demonstrated low-level accelerator-kernel engineering in CUDA, Triton, or similar.
  • Fluency with profiler-led performance analysis and the ability to connect measurements to the GPU memory hierarchy and execution model.
  • A rigorous approach to correctness: reference comparisons, numerical-tolerance policies, and regression gates for optimized code.
  • Judgment about when specialization is warranted and how to state a kernel's envelope while retaining reliable fallbacks.
  • Ability to communicate compact, reproducible evidence to collaborators, separating kernel findings from broader conclusions.

Responsibilities

  • Profile representative workloads to identify kernel bottlenecks and execution regimes.
  • Design, implement, and tune priority kernels in a low-level accelerator stack (CUDA, Triton, or equivalent).
  • Use hardware counters and experiments to reason about memory traffic, cache behavior, and occupancy.
  • Maintain numerical-correctness coverage across shapes, dtypes, layouts, and boundary conditions.
  • Maintain versioned microbenchmarks and regression reports documenting environment and baselines.
  • Document kernel paths, constraints, and safe fallbacks for compiler/runtime partners.

Skills

Low-level kernel engineering
CUDA
Triton
Profiling & analysis
Numerical correctness
Regression testing
Communication of results

Tools

CUDA
Triton

Job description

Make critical operators fast, portable where possible, and correct by construction and test.

Status Open

Area Kernels

Build the critical operator paths that make General Diffusion’s heterogeneous-compute research measurable in real execution. You will turn profiler evidence into carefully tuned kernels, then make the performance claim inseparable from numerical-correctness and regression evidence. The work advances the company’s silicon-neutral direction by making target-specific assumptions explicit and preserving portability wherever measurements support it.

01 / The work
What you’ll work on
  • Profile representative workloads to identify the operators and execution regimes where kernel work can materially affect end-to-end performance; distinguish a kernel bottleneck from a compiler, runtime, or fabric bottleneck before optimizing.
  • Design, implement, and tune priority kernels in an appropriate low-level accelerator stack (for example CUDA, Triton, or an equivalent), with deliberate choices around data layout, tiling, launch configuration, and fusion.
  • Use hardware counters and repeatable experiments to reason about memory traffic, cache behavior, register pressure, occupancy, instruction mix, and synchronization rather than optimizing from timing alone.
  • Build and maintain numerical-correctness coverage against clear reference implementations across meaningful shapes, dtypes, layouts, boundary conditions, and tolerance policies.
  • Maintain versioned microbenchmarks and workload-level regressions that record the environment, inputs, baselines, latency or throughput distributions, and known trade-offs behind each optimization.
  • Document each target’s supported kernel paths, constraints, and safe fallbacks, and provide compiler and runtime partners the operator-level capability information they need without owning compiler-wide lowering or production placement.
02 / The background
What you bring
  • Demonstrated low-level accelerator-kernel engineering in CUDA, Triton, or a comparable environment, including work on performance-sensitive tensor, reduction, normalization, attention, or matrix-multiplication-style operators.
  • Fluency with profiler-led performance analysis and the ability to connect measurements to the GPU memory hierarchy, parallel execution model, and instruction-level behavior.
  • A rigorous approach to correctness: experience constructing reference comparisons, numerical-tolerance policies, edge-case tests, and regression gates for optimized code.
  • Judgment about when specialization is warranted, how to state a kernel’s supported envelope, and how to retain a reliable fallback when assumptions do not hold.
  • Ability to communicate compact, reproducible evidence to compiler, runtime, research, and verification collaborators, separating an operator-level finding from a broader systems conclusion.
03 / The evidence
What progress looks like
  • A reproducible priority-kernel corpus exists for representative workloads, pairing baseline and optimized results with recorded software/hardware context, profiler evidence, and a clear statement of the applicable input regime.
  • Critical kernel changes are covered by automated reference comparisons and numerical-regression cases that exercise relevant shapes, layouts, dtypes, and boundary conditions; failures identify the affected configuration rather than merely reporting a generic mismatch.
  • Compiler and runtime partners can consume an up-to-date, evidence-backed description of kernel capabilities, limits, and fallback conditions for supported targets, making operator choices auditable without transferring ownership of lowering, placement, or fleet behavior.
04 / In the system
Where this role fits

Owns kernel hot paths, not compiler-wide lowering or multi-node fleet orchestration.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, GPU Kernels
Member of Technical Staff, GPU Kernels

SF Tensor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, ML Compilers & Code Generation
Member of Technical Staff, ML Compilers & Code Generation

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Kernel Engineer
Kernel Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff, GPU Kernels
Member of Technical Staff, GPU Kernels

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Equity
Kernel Performance Engineer (CUDA/Triton)
Kernel Performance Engineer (CUDA/Triton)

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Member of Technical Staff, GPU Systems & Fabric
Member of Technical Staff, GPU Systems & Fabric

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Member of Technical Staff, Heterogeneous Runtime & Placement
Member of Technical Staff, Heterogeneous Runtime & Placement

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 250,000
Member of Technical Staff, GPU & ASIC Performance Modeling
Member of Technical Staff, GPU & ASIC Performance Modeling

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 290,000
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 150,000 - 350,000