GPU Kernel Engineer for Fast AI Training & Inference

River AI

Palo Alto (CA)

On-site

USD 200,000 - 420,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Visa sponsorship
Relocation assistance
Comprehensive health insurance

Job summary

River AI in Palo Alto, California is seeking exceptional GPU kernel engineers to accelerate large-model training and inference. You will own performance-critical operations, including attention, matrix multiplication, and low-precision compute, and collaborate with researchers and systems engineers to push the speed and efficiency of our stack.

You will design fast kernels, optimize memory and tiling, develop FP8/FP4 mixed-precision work, and validate correctness through rigorous benchmarks.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • Experience optimizing GPU kernels with CUDA, Triton, CUTLASS, CuTe, or comparable tools.
  • Strong understanding of GPU architecture, memory hierarchies, and parallel execution.
  • Proficiency in C++ and Python.
  • Strong foundations in linear algebra, floating-point arithmetic, and numerical computing.
  • Strong debugging and profiling skills, with a collaborative approach to engineering.

Responsibilities

  • Build fast GPU kernels for attention, matrix multiplication, expert routing, and related operations.
  • Optimize memory access, tiling, and synchronization to make efficient use of GPU hardware.
  • Develop FP8, FP4, and mixed-precision kernels while preserving numerical correctness.
  • Accelerate fine-tuning and RL through optimized adapters, backward passes, and fused operations.
  • Profile real workloads and integrate improvements into training and inference runtimes.
  • Build reproducible benchmarks that verify correctness, gradients, and performance.

Skills

CUDA
GPU architecture
Python
C++

Education

Bachelor's degree in Computer Science/Engineering

Tools

Triton
CUTLASS
CuTe
PyTorch

Job description

River AI in Palo Alto, California is seeking exceptional GPU kernel engineers to accelerate large-model training and inference. You will own performance-critical operations, including attention, matrix multiplication, and low-precision compute, and collaborate with researchers and systems engineers to push the speed and efficiency of our stack.

You will design fast kernels, optimize memory and tiling, develop FP8/FP4 mixed-precision work, and validate correctness through rigorous benchmarks.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Kernel Engineer — Accelerate AI Training & Inference
GPU Kernel Engineer — Accelerate AI Training & Inference

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Software Engineer, GPU Kernels
Software Engineer, GPU Kernels

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Equity
Visa sponsorship
Relocation assistance
+1
Software Engineer, GPU Kernels
Software Engineer, GPU Kernels

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
GPU Kernel Engineer: Build Fast AI Inference at Scale
GPU Kernel Engineer: Build Fast AI Inference at Scale

Baseten • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
100% medical coverage
Generous PTO policy
+2
GPU Kernel Engineer for High-Performance AI Inference
GPU Kernel Engineer for High-Performance AI Inference

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
Systems Engineer - High-Performance AI Training
Systems Engineer - High-Performance AI Training

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Distributed Training Engineer — High-Perf GPU Scale
Distributed Training Engineer — High-Perf GPU Scale

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Relocation assistance
Visa sponsorship
Performance Kernel Engineer for Custom AI Silicon
Performance Kernel Engineer for Custom AI Silicon

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health, dental, and vision benefits
Unlimited PTO
Relocation support
GPU Kernel Engineer for AI Inference & Performance
GPU Kernel Engineer for AI Inference & Performance

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Inference Systems Engineer — High-Performance AI Serving
Inference Systems Engineer — High-Performance AI Serving

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3