GPU Kernel Engineer — Accelerate AI Training & Inference

River AI Inc.

Palo Alto (CA)

On-site

USD 200,000 - 420,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Health insurance
Dental insurance
Vision insurance
Unlimited PTO
Relocation assistance
Visa sponsorship

Job summary

River AI Inc. is seeking exceptional GPU kernel engineers to accelerate training and inference for large models, building compute primitives, attention, matmul, and low-precision kernels.

You’ll own performance-critical ops, optimize memory, validate correctness, and collaborate with researchers and systems engineers to bring improvements into production, enabling faster, more efficient AI workloads.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • Experience optimizing GPU kernels with CUDA, Triton, CUTLASS, CuTe, or comparable tools.
  • Strong understanding of GPU architecture, memory hierarchies, and parallel execution.
  • Proficiency in C++ and Python.
  • Strong foundations in linear algebra, floating-point arithmetic, and numerical computing.
  • Strong debugging and profiling skills, with a collaborative approach to engineering.

Responsibilities

  • Build fast GPU kernels for attention, matrix multiplication, expert routing, and related operations.
  • Optimize memory access, tiling, and synchronization to use GPU hardware efficiently.
  • Develop FP8, FP4, and mixed-precision kernels while preserving numerical correctness.
  • Accelerate fine-tuning and RL through optimized adapters, backward passes, and fused operations.
  • Profile real workloads and integrate improvements into training and inference runtimes.
  • Build reproducible benchmarks that verify correctness, gradients, and performance.

Skills

GPU kernel optimization
C++ programming
Python programming
Profiling & debugging
Numerical computing

Education

Bachelor's degree in CS/CE or equivalent

Tools

CUDA
Triton
CUTLASS
CuTe
PyTorch
vLLM

Job description

River AI Inc. is seeking exceptional GPU kernel engineers to accelerate training and inference for large models, building compute primitives, attention, matmul, and low-precision kernels.

You’ll own performance-critical ops, optimize memory, validate correctness, and collaborate with researchers and systems engineers to bring improvements into production, enabling faster, more efficient AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, GPU Kernels
Software Engineer, GPU Kernels

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
GPU Kernel Engineer: Build Fast AI Inference at Scale
GPU Kernel Engineer: Build Fast AI Inference at Scale

Baseten • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation
100% medical coverage
Generous PTO policy
+2
Inference Systems Engineer - Fast, Multi-GPU Model Serving
Inference Systems Engineer - Fast, Multi-GPU Model Serving

River AI Inc. • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Comprehensive health, dental, and vis
Unlimited PTO
Relocation assistance
+1
GPU Kernel Engineer for High-Performance AI Inference
GPU Kernel Engineer for High-Performance AI Inference

Baseten • San Francisco (CA)

On-site
USD 180,000 - 360,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
+3
GPU Kernel Engineer for AI Inference & Performance
GPU Kernel Engineer for AI Inference & Performance

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
GPU Kernel Architect for High-Performance AI Inference
GPU Kernel Architect for High-Performance AI Inference

Bitdeer (NASDAQ: BTDR) • Austin (TX)

On-site
USD 120,000 - 180,000
Software Engineer – GPU Kernel
Software Engineer – GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
GPU Kernel Engineer — High-Performance ML at Scale
GPU Kernel Engineer — High-Performance ML at Scale

The Consensus • San Francisco (CA)

On-site
USD 120,000 - 160,000
Competitive compensation with equity
100% medical, dental, and vision insurance coverage
Flexible PTO policy including a Winter Break
+2
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000