Stand out for this role — generate a tailored resume and cover letter in about a minute.
Sail builds the world’s most efficient software for inference and agent hosting. In this role, you’ll own token processing down to the lowest layers of the stack, optimize kernel performance, develop new request scheduling and parallelism strategies, and help us use a heterogeneous mix of hardware at max efficiency.
You’ll design and implement exotic parallelism schemes, write custom GPU kernels for regimes like cascade attention, and understand every microsecond of GPU time spent during a
Sail builds the world's most efficient software for inference (processing LLM tokens) and agent hosting (cloud VMs). Together, our technologies allow our customers to deploy AI agents at large scale to do the most challenging work.
In this role, you'll own token processing down to the lowest layers of the stack. You'll optimize kernel performance, develop new request scheduling and parallelism strategies at the engine level, and help us use a heterogenous mix of hardware at max efficiency.