Software Engineer, GPU Kernels

River AI

Palo Alto (CA)

On-site

USD 200,000 - 420,000

Full time

6 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Visa sponsorship
Relocation assistance
Comprehensive health insurance

Job summary

River AI in Palo Alto, California is seeking exceptional GPU kernel engineers to accelerate large-model training and inference. You will own performance-critical operations, including attention, matrix multiplication, and low-precision compute, and collaborate with researchers and systems engineers to push the speed and efficiency of our stack.

You will design fast kernels, optimize memory and tiling, develop FP8/FP4 mixed-precision work, and validate correctness through rigorous benchmarks.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • Experience optimizing GPU kernels with CUDA, Triton, CUTLASS, CuTe, or comparable tools.
  • Strong understanding of GPU architecture, memory hierarchies, and parallel execution.
  • Proficiency in C++ and Python.
  • Strong foundations in linear algebra, floating-point arithmetic, and numerical computing.
  • Strong debugging and profiling skills, with a collaborative approach to engineering.

Responsibilities

  • Build fast GPU kernels for attention, matrix multiplication, expert routing, and related operations.
  • Optimize memory access, tiling, and synchronization to make efficient use of GPU hardware.
  • Develop FP8, FP4, and mixed-precision kernels while preserving numerical correctness.
  • Accelerate fine-tuning and RL through optimized adapters, backward passes, and fused operations.
  • Profile real workloads and integrate improvements into training and inference runtimes.
  • Build reproducible benchmarks that verify correctness, gradients, and performance.

Skills

CUDA
GPU architecture
Python
C++

Education

Bachelor's degree in Computer Science/Engineering

Tools

Triton
CUTLASS
CuTe
PyTorch

Job description

At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.

Who we are

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

About The Role

We are looking for exceptional GPU kernel engineers to build the compute primitives behind River’s training and inference infrastructure. Your goal is to make large models faster to train and more efficient to serve.

You will own performance-critical operations, including attention, matrix multiplication, mixture-of-experts execution, and low-precision computation. Working closely with researchers and systems engineers, you will identify bottlenecks, implement kernels, validate correctness, and bring improvements into production.

What You’ll Do
  • Build fast GPU kernels for attention, matrix multiplication, expert routing, and related operations.
  • Optimize memory access, tiling, and synchronization to make efficient use of GPU hardware.
  • Develop FP8, FP4, and mixed-precision kernels while preserving numerical correctness.
  • Accelerate fine-tuning and RL through optimized adapters, backward passes, and fused operations.
  • Profile real workloads and integrate improvements into training and inference runtimes.
  • Build reproducible benchmarks that verify correctness, gradients, and performance.
Skills & Qualifications
Minimum Qualifications:
  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • Experience optimizing GPU kernels with CUDA, Triton, CUTLASS, CuTe, or comparable tools.
  • Strong understanding of GPU architecture, memory hierarchies, and parallel execution.
  • Proficiency in C++ and Python.
  • Strong foundations in linear algebra, floating-point arithmetic, and numerical computing.
  • Strong debugging and profiling skills, with a collaborative approach to engineering.
Preferred Qualifications:

(We encourage you to apply even if you don't meet all of these)

  • Experience optimizing for NVIDIA Blackwell or Hopper GPUs.
  • Work on attention, mixture-of-experts kernels, grouped GEMMs, or low-rank adapters.
  • Experience implementing backward passes and validating gradients.
  • Familiarity with FP8, FP4, and quantized weight layouts.
  • Experience integrating custom operators into PyTorch, SGLang, vLLM, or similar frameworks.
  • Open-source contributions or a track record of shipping substantial kernel optimizations.
Logistics & Benefits
  • Location: Palo Alto, California.
  • Compensation: Depending on experience and skills the expected base pay is $200,000 - $420,000 USD per year, plus equity.
  • Benefits: Comprehensive health, dental, and vision insurance; unlimited PTO; and relocation assistance as needed.
  • Visa Sponsorship: We sponsor visas and are committed to supporting the process for the right candidate.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, GPU Kernels
Software Engineer, GPU Kernels

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Software Engineer, River API
Software Engineer, River API

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Kernel Engineer (Custom Silicon), Hardware
Kernel Engineer (Custom Silicon), Hardware

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision benefits
Unlimited PTO
Relocation support
Software Engineer – GPU Kernel
Software Engineer – GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Member of Technical Staff, GPU Kernels
Member of Technical Staff, GPU Kernels

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 285,000 - 315,000
Relocation assistance
Equity
Software Engineer, Inference Systems
Software Engineer, Inference Systems

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Comprehensive health, dental, and vis
Unlimited PTO
Relocation assistance
+1
GPU Kernel Engineer — Accelerate AI Training & Inference
GPU Kernel Engineer — Accelerate AI Training & Inference

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Member of Technical Staff - GPU Performance Engineer
Member of Technical Staff - GPU Performance Engineer

Liquid AI • San Francisco (CA)

On-site
USD 120,000 - 180,000
Competitive base salary with equity
100% medical, dental, and vision premiums
401(k) matching up to 4%
+2
Member of Technical Staff, GPU Kernels
Member of Technical Staff, GPU Kernels

SF Tensor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - GPU Performance Engineer San Francisco · Remote · Boston · Hybrid
Member of Technical Staff - GPU Performance Engineer San Francisco · Remote · Boston · Hybrid

Liquid AI, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Equity in unicorn-stage company
Health premiums paid
401(k) matching up to 4%
+2