Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
River AI in Palo Alto is seeking exceptional inference systems engineers to build the engines that serve large models through the River API. Your goal is to deliver fast, reliable inference while making efficient use of GPU compute and memory.
You will own the serving runtime from request scheduling and continuous batching to KV-cache management, distributed model execution, and checkpoint loading. Working closely with GPU kernel engineers, researchers, and infrastructure engineers, you will
River AI in Palo Alto is seeking exceptional inference systems engineers to build the engines that serve large models through the River API. Your goal is to deliver fast, reliable inference while making efficient use of GPU compute and memory.
You will own the serving runtime from request scheduling and continuous batching to KV-cache management, distributed model execution, and checkpoint loading. Working closely with GPU kernel engineers, researchers, and infrastructure engineers, you will