Kernel Performance Engineer (CUDA/Triton)

General Diffusion, Inc.

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 280,000

Full time

11 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

General Diffusion, Inc. in San Francisco seeks a senior engineer to own kernel hot paths across heterogeneous accelerators, turning profiler evidence into portable, numerically correct kernels.

You will profile workloads, design and tune kernels in CUDA/Triton, optimize data layouts and launch configurations, and maintain regression tests and reference implementations to prove performance and correctness. The role emphasizes portability, clear documentation for compiler and runtime partners, and

Qualifications

  • Demonstrated low-level accelerator-kernel engineering in CUDA, Triton, or similar.
  • Fluency with profiler-led performance analysis and the ability to connect measurements to the GPU memory hierarchy and execution model.
  • A rigorous approach to correctness: reference comparisons, numerical-tolerance policies, and regression gates for optimized code.
  • Judgment about when specialization is warranted and how to state a kernel's envelope while retaining reliable fallbacks.
  • Ability to communicate compact, reproducible evidence to collaborators, separating kernel findings from broader conclusions.

Responsibilities

  • Profile representative workloads to identify kernel bottlenecks and execution regimes.
  • Design, implement, and tune priority kernels in a low-level accelerator stack (CUDA, Triton, or equivalent).
  • Use hardware counters and experiments to reason about memory traffic, cache behavior, and occupancy.
  • Maintain numerical-correctness coverage across shapes, dtypes, layouts, and boundary conditions.
  • Maintain versioned microbenchmarks and regression reports documenting environment and baselines.
  • Document kernel paths, constraints, and safe fallbacks for compiler/runtime partners.

Skills

Low-level kernel engineering
CUDA
Triton
Profiling & analysis
Numerical correctness
Regression testing
Communication of results

Tools

CUDA
Triton

Job description

General Diffusion, Inc. in San Francisco seeks a senior engineer to own kernel hot paths across heterogeneous accelerators, turning profiler evidence into portable, numerically correct kernels.

You will profile workloads, design and tune kernels in CUDA/Triton, optimize data layouts and launch configurations, and maintain regression tests and reference implementations to prove performance and correctness. The role emphasizes portability, clear documentation for compiler and runtime partners, and

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Kernel Engineer: 26-02794
Senior Kernel Engineer: 26-02794

Akraya, Inc. • Bellevue (WA)

On-site
USD 117,000 - 124,000
Member of Technical Staff, Kernels
Member of Technical Staff, Kernels

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000
Inference Runtime Performance Engineer — GPU Kernels
Inference Runtime Performance Engineer — GPU Kernels

iFrame Corporation • San Francisco (CA)

Remote
USD 220,000 - 360,000
CUDA Kernel Engineer — Optimize GPU Performance at Scale
CUDA Kernel Engineer — Optimize GPU Performance at Scale

Pragmatike • California (MO)

On-site
USD 180,000 - 240,000
Health, Dental, and Vision
Sign-on bonus
401k
Kernel Engineer for High-Performance ML Compute (CUDA/Triton)
Kernel Engineer for High-Performance ML Compute (CUDA/Triton)

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior Triton GPU Compiler & Kernel Engineer
Senior Triton GPU Compiler & Kernel Engineer

AMD • San Jose (CA)

Hybrid
USD 180,000 - 240,000
AMD Benefits
Staff Engineer, GPU Systems & Fabric
Staff Engineer, GPU Systems & Fabric

General Diffusion, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 250,000
Kernel Engineer
Kernel Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Kernel Engineer – CUDA/Triton Optimizer
Senior AI Kernel Engineer – CUDA/Triton Optimizer

Akraya, Inc. • Bellevue (WA)

On-site
USD 117,000 - 124,000
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000