Stand out for this role — generate a tailored resume and cover letter in about a minute.
Sail Research in San Francisco is looking for a motivated software engineer to optimize token processing at every layer of the stack. You will modify inference engines and analyze GPU performance, ensuring efficient hardware utilization.
The ideal candidate has a solid understanding of LLM mechanics and interests in cutting-edge MLSys research. The company offers benefits like free meals and a Studio Display for every employee.
Optimize token processing down to the lowest layers of the stack. You'll optimize kernel performance, develop new scheduling and parallelism strategies, and help us squeeze every FLOP out of our hardware.
Modify and extend state-of-the-art inference engines like vLLM and SGLang.
Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain every kernel launch on an NSys profile.
Design and implement exotic parallelism schemes to work with "interesting" hardware topologies.
Write custom GPU kernels to excel in specific regimes, such as cascade attention.
Strong understanding of LLM mechanics, like KV cache, mixture-of-experts, prefill vs. decode phases.
Interest in MLSys research—great ideas like speculative decoding and sparse attention come from research, that we need to follow closely.
Familiarity with modern, tile-based GPU programming, e.g. Triton, CUTLASS, ThunderKittens, etc. Or an interest in learning these!
Meals are provided. Every employee receives a Studio Display.