GPU Kernel Engineer — Kernel Optimization & Compiler Innovation

SF Tensor

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

47 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

SF Tensor is building the fastest GPU compiler and a versatile kernel optimizer. We seek a Member of Technical Staff for GPU Kernel Engineering to push the envelope on what hardware can do before any search runs.

You will write and hand-optimize kernels that execute during pre-training, post-training and inference across NVIDIA, AMD, TPU and Trainium. You will convert kernel learnings into compiler-search structures, with access to an end-to-end stack down to the ISA.

Qualifications

  • Track record of hand-writing kernels that beat vendor libraries or are on par with them.
  • Strong C++/CUDA or ROCm/HIP low-level programming skills.
  • Experience reading GPU ISA and machine-level assembly (PTX/SASS/GCN).
  • Familiarity with profiling tools to diagnose microarchitectural bottlenecks.

Responsibilities

  • Write and hand-optimize kernels for real workloads to maximize ceiling before search space exploration.
  • Profile at microarchitectural level: SM/CU utilization, warp stalls, memory conflicts, register pressure.
  • Debug down to clock behavior, thermal throttling and driver paths; ensure correctness with formal proof where applicable.
  • Convert hard-won kernel knowledge into structures the compiler can search.
  • Work below PTX at the ISA level, reasoning about SASS and cubins for schedules.
  • Build performance models, microbenchmarks and tooling to predict kernel behavior.
  • Collaborate with the formal correctness team to ship kernels with a formal proof when required.

Skills

Hand-tuned kernels
Low-level GPU programming
C++
CUDA
ISA familiarity
Performance profiling

Tools

Nsight Compute
Nsight Systems
rocprof
omniperf

Job description

SF Tensor is building the fastest GPU compiler and a versatile kernel optimizer. We seek a Member of Technical Staff for GPU Kernel Engineering to push the envelope on what hardware can do before any search runs.

You will write and hand-optimize kernels that execute during pre-training, post-training and inference across NVIDIA, AMD, TPU and Trainium. You will convert kernel learnings into compiler-search structures, with access to an end-to-end stack down to the ISA.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Kernel Engineer, High-Performance Compiler
Senior GPU Kernel Engineer, High-Performance Compiler

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Equity
Member of Technical Staff, GPU Kernels
Member of Technical Staff, GPU Kernels

SF Tensor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, GPU Kernels
Member of Technical Staff, GPU Kernels

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Equity
Staff Engineer — AI-Driven GPU Compiler Optimization
Staff Engineer — AI-Driven GPU Compiler Optimization

SF Tensor • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Equity and benefits
Office in San Francisco
Senior Kernel & Compiler Performance Engineer (GPU/AI)
Senior Kernel & Compiler Performance Engineer (GPU/AI)

RadixArk • Palo Alto (CA)

On-site
USD 210,000 - 290,000
Competitive compensation
Comprehensive benefits
Flexible work arrangements
Member of Technical Staff, GPU Compiler
Member of Technical Staff, GPU Compiler

SF Tensor • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Member of Technical Staff, GPU Compiler
Member of Technical Staff, GPU Compiler

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Office in San Francisco
Senior GPU Compiler Engineer (MLIR/LLVM)
Senior GPU Compiler Engineer (MLIR/LLVM)

SF Tensor • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Staff Engineer: GPU Kernels & AI Performance
Staff Engineer: GPU Kernels & AI Performance

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior GPU Compiler Engineer — End-to-End MLIR & CUDA
Senior GPU Compiler Engineer — End-to-End MLIR & CUDA

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Office in San Francisco