Staff AI Inference Kernel Engineer

Sail

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Sail builds the world’s most efficient software for inference and agent hosting. In this role, you’ll own token processing down to the lowest layers of the stack, optimize kernel performance, develop new request scheduling and parallelism strategies, and help us use a heterogeneous mix of hardware at max efficiency.

You’ll design and implement exotic parallelism schemes, write custom GPU kernels for regimes like cascade attention, and understand every microsecond of GPU time spent during a

Qualifications

  • Strong understanding of core LLM mechanics (KV cache, mixture-of-experts) and deployment stages.
  • Interest in MLSys research (speculative decoding, sparse attention).
  • Familiarity with tile-based GPU programming or willingness to learn (Triton, CUTLASS).
  • Excellent communication and ability to avoid relying on generated prose.

Responsibilities

  • Modify and extend state-of-the-art inference engines like vLLM and SGLang.
  • Analyze GPU time per forward pass and explain kernel launches.
  • Design exotic parallelism schemes for diverse hardware topologies.
  • Write custom GPU kernels to optimize regimes like cascade attention.

Skills

LLM basics
GPU programming
System architecture
Communication

Tools

vLLM
SGLang
Triton
CUTLASS

Job description

Sail builds the world’s most efficient software for inference and agent hosting. In this role, you’ll own token processing down to the lowest layers of the stack, optimize kernel performance, develop new request scheduling and parallelism strategies, and help us use a heterogeneous mix of hardware at max efficiency.

You’ll design and implement exotic parallelism schemes, write custom GPU kernels for regimes like cascade attention, and understand every microsecond of GPU time spent during a

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer, AI Inference & Distributed Systems
Staff Engineer, AI Inference & Distributed Systems

Sail Research • San Francisco (CA)

On-site
USD 120,000 - 160,000
Free meals
Studio Display at desk
Friendly office environment with a cat
Staff Engineer — AI Inference & Distributed Systems
Staff Engineer — AI Inference & Distributed Systems

Sail • San Francisco (CA)

On-site
USD 150,000 - 190,000
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Sail • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior AI Inference & Kernel Engineer
Senior AI Inference & Kernel Engineer

Intel • Austin (TX)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Vacation
Senior AI Inference Systems Engineer | GPU Kernels & Runtime
Senior AI Inference Systems Engineer | GPU Kernels & Runtime

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Hybrid AI Inference Engineer — Kernel & Performance
Hybrid AI Inference Engineer — Kernel & Performance

Intel • Hillsboro (OR)

Hybrid
USD 189,000 - 315,000
Stock bonuses
Health benefits
Retirement plan
+1
Staff AI Infrastructure Engineer — Orchestration & Inference
Staff AI Infrastructure Engineer — Orchestration & Inference

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 213,000 - 288,000
Early-stage equity
Direct access to leadership
Senior GPU Kernel Engineer for High-Performance AI Inference
Senior GPU Kernel Engineer for High-Performance AI Inference

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
ESPP
+2
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity
Benefits
Senior GPU Kernel Architect & Optimizer
Senior GPU Kernel Architect & Optimizer

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical–dental–vision insurance
401(k) with employer match
Flexible PTO
+3