Founding GPU Kernel Engineer

SF Tensor

San Francisco (CA)

On-site

USD 285,000 - 315,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

SF Tensor is looking for a Founding GPU Kernel Engineer in San Francisco, specializing in GPU architecture and kernel optimization for machine learning workloads. The ideal candidate has deep expertise, proven capabilities in hand-optimizing performance-critical kernels, and strong programming skills in C++ and CUDA. This full-time position offers a competitive salary of $285,000 - $315,000, plus bonus and equity, with relocation assistance available for the right candidate who values in-person collaboration.

Qualifications

  • Deep expertise in GPU architecture.
  • Proven track record of hand-writing kernels that match or beat vendor libraries.
  • Experience reading and reasoning about PTX/SASS or GPU assembly.

Responsibilities

  • Write and hand-optimize GPU kernels for ML workloads.
  • Profile microarchitectural performance.
  • Debug performance issues related to hardware execution.

Skills

GPU architecture expertise
C++ programming
CUDA programming
Low-level profiling tools
Distributed training systems

Tools

Nsight Compute
Nsight Systems
rocprof

Job description

About The Role

We're looking for a Founding GPU Kernel Engineer who lives right at the boundary between hardware and software. Someone who thinks in warps, occupancy, and memory hierarchies, and can squeeze every last FLOP out of a GPU.

Your job is to go deeper than anyone else. You'll hand-tune kernels to figure out what's actually possible on the hardware, and then turn that knowledge into compiler optimization passes that help every model we compile.

What You'll Do
  • Write and hand-optimize GPU kernels for ML workloads (matmuls, attention, normalization, etc.) to set the performance ceilings
  • Profile at the microarchitectural level: look into SM utilization, warp stalls, memory bank conflicts, register pressure, instruction throughput
  • Debug performance issues by digging deep into things like clock speeds, thermal throttling, driver behavior, hardware errata
  • Turn your hand-optimization insights into automated compiler passes (working closely with our compiler team)
  • Develop performance models that predict how kernels will behave across different GPU architectures
  • Build tools and methods for systematic kernel optimization
  • Work with NVIDIA, AMD, and emerging AI accelerators – understand the common parts and what's vendor-specific
What We're Looking For
  • Deep expertise in GPU architecture
  • Proven track record of hand-writing kernels that match or beat vendor libraries (cuBLAS, cuDNN, CUTLASS)
  • Strong skills with low-level profiling tools: Nsight Compute, Nsight Systems, rocprof, or equivalents
  • Experience reading and reasoning about PTX/SASS or GPU assembly
  • Solid systems programming in C++ and CUDA (or ROCm/HIP)
  • Good understanding of how high-level ML operations map to hardware execution
  • Experience with distributed training systems: collective ops like all-reduce and all-gather, NCCL/RCCL, multi-node communication patterns
Nice to Have
  • HPC background: experience with large-scale scientific computing, MPI, or work in supercomputing
  • Background in electrical engineering, computer architecture, or hardware design
  • Driver development experience (NVIDIA, AMD, or other accelerators)
  • Experience with MLIR, LLVM, or compiler backends
  • Deep knowledge of distributed ML training: gradient accumulation, activation checkpointing, pipeline/tensor parallelism, ZeRO-style optimizations
  • Familiarity with custom accelerators: TPUs, Trainium, Inferentia, or similar
  • Knowledge of high-speed interconnects: NVLink, NVSwitch, InfiniBand, RoCE
  • Publications or contributions in GPU optimization, HPC, or ML systems
  • Experience at NVIDIA, AMD, a national lab, or an AI hardware/infrastructure company
Why Join Us

This role is for someone who wants to know why things are fast or slow on the hardware. You'll have a direct impact on the performance of large-scale AI training, tackling problems that need real depth. If you've ever been annoyed that your hard-won optimization knowledge is stuck in your head and not baked into a compiler, here's your shot to change that.

We believe in the power of in-person collaboration to solve the hardest problems and foster a strong team culture. We offer relocation assistance and look forward to you joining us in our San Francisco office.

Salary & Benefits

The base salary range for this full-time position is $285,000 - $315,000 + bonus + equity + benefits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding GPU Kernel Engineer
Founding GPU Kernel Engineer

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Equity
Comprehensive benefits package
Founding GPU Kernel Engineer — Hand‑Tuned ML Kernels
Founding GPU Kernel Engineer — Hand‑Tuned ML Kernels

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Equity
Comprehensive benefits package
Software Engineer – GPU Kernel
Software Engineer – GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
KERNEL ENGINEER
KERNEL ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff - GPU Performance Engineer
Member of Technical Staff - GPU Performance Engineer

Liquid AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Competitive base salary with equity
100% medical, dental, and vision premiums
401(k) matching up to 4%
+2
GPU Performance / Kernel Engineer
GPU Performance / Kernel Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan
Annual bonus
+1
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 150,000 - 350,000
Pioneering GPU Kernel Engineer for ML Performance
Pioneering GPU Kernel Engineer for ML Performance

SF Tensor • San Francisco (CA)

On-site
USD 285,000 - 315,000
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

NVIDIA • Austin (TX)

On-site
USD 184,000
Equity
Benefits