KERNEL ENGINEER

MakerMaker.AI

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

MakerMaker.AI in San Francisco is seeking a skilled Software Engineer to write and optimize GPU kernels. You will work on deep low-level tasks that directly impact the performance of machine learning models.

The ideal candidate has over 4 years of experience with GPU kernels, strong systems expertise, and a proven track record in kernel optimizations. This role requires on-site work in a collaborative environment.

Qualifications

  • 4+ years writing performant GPU kernels (CUDA, ROCm, Triton, or equivalent).
  • Fluency with memory hierarchy, occupancy, and tensor cores.
  • Previous experience shipping kernel-level optimizations.

Responsibilities

  • Write and optimize GPU kernels for training and inference workloads.
  • Profile workloads and translate findings into kernel optimizations.
  • Integrate optimized kernels into training and serving stacks.

Skills

Performance optimization
CUDA
Profiling tools
Systems expertise
Python
C++

Tools

Nsight
ncu

Job description

ABOUT THE COMPANY

We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site

ABOUT THE ROLE

You’ll write and optimize the GPU kernels and supporting systems software that makes our training and inference workloads fast. This is deep, low-level work (performance counters, memory bandwidth, warp-level scheduling) applied to the specific shapes and patterns our models actually use.

We hire kernel engineers because the gap between "this works" and "this is fast on the hardware we have" is enormous, and that gap directly bounds what our researchers can try. You’ll close that gap.

WHAT YOU'LL DO
  • Write and optimize GPU kernels (CUDA, ROCm, Triton, or similar) for training and inference workloads: attention variants, MoE layers, custom activations, communication primitives
  • Profile real workloads with hardware counters and translate findings into specific kernel-level optimizations
  • Co‑design kernels with the research teams, when the kernel and the algorithm need to change together, you participate in both
  • Integrate optimized kernels into our training and serving stacks; benchmark before and after; verify the win is real end‑to‑end
  • Maintain kernel quality over time as hardware, frameworks, and workloads shift underneath
  • Spread kernel‑level fluency across the team; we want this expertise shared, not siloed
WHAT WE'RE LOOKING FOR
  • 4+ years writing performant GPU kernels (CUDA, ROCm, Triton, or production‑grade equivalent)
  • Hardware‑level fluency: memory hierarchy, occupancy, register pressure, tensor cores, warp scheduling
  • Profiling fluency (Nsight, ncu, or comparable tools) and the discipline to measure before changing
  • Track record of shipping kernel‑level optimizations that moved a measurable metric in a real system
  • Strong systems expertise: you understand how kernels live inside larger frameworks and how integration choices affect end‑to‑end performance
  • Comfortable reading framework‑level Python and C++ around your kernels
NICE TO HAVE
  • Open‑source contributions to kernel libraries, compilers, or ML frameworks
  • Experience with multiple accelerator architectures (different GPU families, TPUs, custom ASICs), preferably AMD GPUs
  • Familiarity with collective communication primitives (NCCL or equivalent)
  • Compiler or runtime background
THIS ROLE IS PROBABLY NOT FOR YOU IF
  • You haven’t gotten your hands dirty at the kernel level: this isn’t a higher‑level systems role rebranded
  • You want to stay narrowly in one library; we expect breadth across the kernel surface our models actually use
  • Performance work without measurable end‑to‑end impact frustrates you
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding GPU Kernel Engineer
Founding GPU Kernel Engineer

SF Tensor • San Francisco (CA)

On-site
USD 285,000 - 315,000
CUDA Engineer - Kernel Optimization
CUDA Engineer - Kernel Optimization

Mercor • San Francisco (CA)

On-site
USD 96,000 - 165,000
CUDA Engineer - Kernel Optimization - AI Trainer
CUDA Engineer - Kernel Optimization - AI Trainer

Mercor • Chicago (IL)

On-site
USD 83,000 - 165,000
Founding GPU Kernel Engineer
Founding GPU Kernel Engineer

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Equity
Comprehensive benefits package
Software Engineer – GPU Kernel
Software Engineer – GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Kernel Engineer (Internship and Full-time)
Kernel Engineer (Internship and Full-time)

Tilde Research • Palo Alto (CA)

On-site
USD 150,000 - 260,000
GPU Performance / Kernel Engineer
GPU Performance / Kernel Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan
Annual bonus
+1
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Bala Cynwyd (PA)

On-site
USD 120,000 - 160,000
Member of Technical Staff - Kernels & GPU Performance
Member of Technical Staff - Kernels & GPU Performance

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000