A complete application in a minute — tailored resume and cover letter, ready to send.
River AI in Palo Alto, California is seeking exceptional GPU kernel engineers to accelerate large-model training and inference. You will own performance-critical operations, including attention, matrix multiplication, and low-precision compute, and collaborate with researchers and systems engineers to push the speed and efficiency of our stack.
You will design fast kernels, optimize memory and tiling, develop FP8/FP4 mixed-precision work, and validate correctness through rigorous benchmarks.
At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.
We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.
We are looking for exceptional GPU kernel engineers to build the compute primitives behind River’s training and inference infrastructure. Your goal is to make large models faster to train and more efficient to serve.
You will own performance-critical operations, including attention, matrix multiplication, mixture-of-experts execution, and low-precision computation. Working closely with researchers and systems engineers, you will identify bottlenecks, implement kernels, validate correctness, and bring improvements into production.
(We encourage you to apply even if you don't meet all of these)