Member of Technical Staff - Inference

Sail

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Sail builds the world’s most efficient software for inference and agent hosting. In this role, you’ll own token processing down to the lowest layers of the stack, optimize kernel performance, develop new request scheduling and parallelism strategies, and help us use a heterogeneous mix of hardware at max efficiency.

You’ll design and implement exotic parallelism schemes, write custom GPU kernels for regimes like cascade attention, and understand every microsecond of GPU time spent during a

Qualifications

  • Strong understanding of core LLM mechanics (KV cache, mixture-of-experts) and deployment stages.
  • Interest in MLSys research (speculative decoding, sparse attention).
  • Familiarity with tile-based GPU programming or willingness to learn (Triton, CUTLASS).
  • Excellent communication and ability to avoid relying on generated prose.

Responsibilities

  • Modify and extend state-of-the-art inference engines like vLLM and SGLang.
  • Analyze GPU time per forward pass and explain kernel launches.
  • Design exotic parallelism schemes for diverse hardware topologies.
  • Write custom GPU kernels to optimize regimes like cascade attention.

Skills

LLM basics
GPU programming
System architecture
Communication

Tools

vLLM
SGLang
Triton
CUTLASS

Job description

Sail builds the world's most efficient software for inference (processing LLM tokens) and agent hosting (cloud VMs). Together, our technologies allow our customers to deploy AI agents at large scale to do the most challenging work.

In this role, you'll own token processing down to the lowest layers of the stack. You'll optimize kernel performance, develop new request scheduling and parallelism strategies at the engine level, and help us use a heterogenous mix of hardware at max efficiency.

What you’ll do
  • Modify and extend state-of-the-art inference engines like vLLM and SGLang.
  • Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain every kernel launch on an NSys profile.
  • Design and implement exotic parallelism schemes to work with "interesting" hardware topologies.
  • Write custom GPU kernels to excel in specific regimes, such as cascade attention.
What we’re looking for
  • Strong understanding of core LLM mechanics, like KV cache, mixture-of-experts, prefill vs. decode phases.
  • Interest in MLSys research - great ideas like speculative decoding and sparse attention come from research, that we need to follow closely.
  • Familiarity with modern, tile-based GPU programming, e.g. Triton, CUTLASS, ThunderKittens, etc. Or an interest in learning these!
  • Great interpersonal communication - please don't use LLMs to write prose. We desk-reject most LLM-generated cover letters and resumes.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Inference
Member of Technical Staff - Inference

Sail Research • San Francisco (CA)

On-site
USD 120,000 - 160,000
Meals provided
Studio Display for every employee
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Member of Technical Staff, MLSys
Member of Technical Staff, MLSys

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Machine Learning Engineer (LLM inference)
Machine Learning Engineer (LLM inference)

GMI Cloud • Mountain View (CA)

On-site
USD 180,000 - 240,000