Founding GPU Kernel Engineer

San Francisco Tensor Company

San Francisco (CA)

On-site

USD 285,000 - 315,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Relocation assistance
Equity
Comprehensive benefits package

Job summary

San Francisco Tensor Company is seeking a Founding GPU Kernel Engineer to enhance GPU performance for AI applications. You will optimize and write kernels while collaborating with compiler teams to improve efficiencies across architectures. The ideal candidate has deep expertise in GPU architecture and experience with profiling tools.

This position offers a salary range of $285,000 - $315,000 along with bonus, equity, and benefits, plus relocation assistance to our San Francisco office.

Qualifications

  • Deep expertise in GPU architecture required.
  • Proven record of hand-writing GPU kernels essential.
  • Strong skills with low-level profiling tools are a must.

Responsibilities

  • Write and hand-optimize GPU kernels for ML workloads.
  • Profile microarchitectural performance issues.
  • Develop performance models for kernels across architectures.

Skills

GPU architecture
C++
CUDA
Low-level profiling tools
ML workloads optimization

Tools

Nsight Compute
Nsight Systems
ROCm/HIP

Job description

About SF Tensor

At The San Francisco Tensor Company, we believe the future of AI and high-performance computing depends on rethinking the entire software and infrastructure stack. Today's developers face bottlenecks across hardware, cloud, and code optimization that slow progress before ideas can reach their full potential. Our mission is to remove those barriers and make compute faster, cheaper, and universally portable.

We are building a Kernel Optimizer that automatically transforms code into its most efficient form, combined with Tensor Cloud for adaptive, cross-cloud compute and Emma Lang, a new programming language for high-performance, hardware‑aware computation. Together, these technologies reinvent the foundations of AI and HPC.

SF Tensor is proudly backed by Susa Ventures and Y Combinator, as well as a group of angels including Max Mullen and Paul Graham and founders and executives of NeuraLink, Notion and AMD. We are partnering with researchers, engineers, and organizations who share our belief that the next breakthroughs in AI require breakthroughs in compute.

About the Role

We're looking for a Founding GPU Kernel Engineer who lives right at the boundary between hardware and software. Someone who thinks in warps, occupancy, and memory hierarchies, and can squeeze every last FLOP out of a GPU.

Your job is to go deeper than anyone else. You'll hand‑tune kernels to figure out what's actually possible on the hardware, and then turn that knowledge into compiler optimization passes that help every model we compile.

What You'll Do
  • Write and hand‑optimize GPU kernels for ML workloads (matmuls, attention, normalization, etc.) to set the performance ceilings
  • Profile at the microarchitectural level: look into SM utilization, warp stalls, memory bank conflicts, register pressure, instruction throughput
  • Debug performance issues by digging deep into things like clock speeds, thermal throttling, driver behavior, hardware errata
  • Turn your hand‑optimization insights into automated compiler passes (working closely with our compiler team)
  • Develop performance models that predict how kernels will behave across different GPU architectures
  • Build tools and methods for systematic kernel optimization
  • Work with NVIDIA, AMD, and emerging AI accelerators - understand the common parts and what's vendor‑specific
What We're Looking For
  • Deep expertise in GPU architecture
  • Proven track record of hand‑writing kernels that match or beat vendor libraries (cuBLAS, cuDNN, CUTLASS)
  • Strong skills with low‑level profiling tools: Nsight Compute, Nsight Systems, rocprof, or equivalents
  • Experience reading and reasoning about PTX/SASS or GPU assembly
  • Solid systems programming in C++ and CUDA (or ROCm/HIP)
  • Good understanding of how high‑level ML operations map to hardware execution
  • Experience with distributed training systems: collective ops like all‑reduce and all‑gather, NCCL/RCCL, multi‑node communication patterns
Nice to Have
  • HPC background: experience with large‑scale scientific computing, MPI, or work in supercomputing
  • Background in electrical engineering, computer architecture, or hardware design
  • Driver development experience (NVIDIA, AMD, or other accelerators)
  • Experience with MLIR, LLVM, or compiler backends
  • Deep knowledge of distributed ML training: gradient accumulation, activation checkpointing, pipeline/tensor parallelism, ZeRO‑style optimizations
  • Familiarity with custom accelerators: TPUs, Trainium, Inferentia, or similar
  • Knowledge of high‑speed interconnects: NVLink, NVSwitch, InfiniBand, RoCE
  • Publications or contributions in GPU optimization, HPC, or ML systems
  • Experience at NVIDIA, AMD, a national lab, or an AI hardware/infrastructure company
Why Join Us

This role is for someone who wants to know why things are fast or slow on the hardware. You'll have a direct impact on the performance of large‑scale AI training, tackling problems that need real depth. If you've ever been annoyed that your hard‑won optimization knowledge is stuck in your head and not baked into a compiler, here's your shot to change that.

We believe in the power of in‑person collaboration to solve the hardest problems and foster a strong team culture. We offer relocation assistance and look forward to you joining us in our San Francisco office.

The base salary range for this full‑time position is $285,000 - $315,000 + bonus + equity + benefits.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding GPU Kernel Engineer
Founding GPU Kernel Engineer

SF Tensor • San Francisco (CA)

On-site
USD 285,000 - 315,000
Founding GPU Compiler Engineer
Founding GPU Compiler Engineer

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Bonus
Equity
Founding GPU Kernel Engineer — Hand‑Tuned ML Kernels
Founding GPU Kernel Engineer — Hand‑Tuned ML Kernels

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Equity
Comprehensive benefits package
Member of Technical Staff - GPU Performance Engineer
Member of Technical Staff - GPU Performance Engineer

Liquid AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Competitive base salary with equity
100% medical, dental, and vision premiums
401(k) matching up to 4%
+2
Software Engineer – GPU Kernel
Software Engineer – GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
GPU Performance / Kernel Engineer
GPU Performance / Kernel Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan
Annual bonus
+1
Founding Product Engineer
Founding Product Engineer

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 225,000 - 275,000
Relocation assistance
Equity
Comprehensive benefits
Founding Research Engineer, AI-Driven Compilation
Founding Research Engineer, AI-Driven Compilation

SF Tensor • San Francisco (CA)

On-site
USD 275,000 - 315,000
Founding Research Engineer, AI-Driven Compilation
Founding Research Engineer, AI-Driven Compilation

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Equity options
Comprehensive benefits
Software Engineer - GPU Kernels
Software Engineer - GPU Kernels

The Consensus • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation with equity
100% medical, dental, and vision insurance coverage
Flexible PTO policy including a Winter Break
+2