Member of Technical Staff - Inference

Sail Research

San Francisco (CA)

On-site

USD 120,000 - 160,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Meals provided
Studio Display for every employee

Job summary

Sail Research in San Francisco is looking for a motivated software engineer to optimize token processing at every layer of the stack. You will modify inference engines and analyze GPU performance, ensuring efficient hardware utilization.

The ideal candidate has a solid understanding of LLM mechanics and interests in cutting-edge MLSys research. The company offers benefits like free meals and a Studio Display for every employee.

Qualifications

  • Strong understanding of LLM mechanics, such as KV cache and mixture-of-experts.
  • Interest in MLSys research and current advancements.
  • Familiarity or willingness to learn modern tile-based GPU programming.

Responsibilities

  • Modify and extend state-of-the-art inference engines.
  • Analyze GPU time utilization and explain kernel launches.
  • Design and implement parallelism schemes for unique hardware.

Skills

Understanding of LLM mechanics
GPU programming
Interest in MLSys research

Tools

Triton
CUTLASS
ThunderKittens

Job description

Optimize token processing down to the lowest layers of the stack. You'll optimize kernel performance, develop new scheduling and parallelism strategies, and help us squeeze every FLOP out of our hardware.

What you’ll do
  • Modify and extend state-of-the-art inference engines like vLLM and SGLang.

  • Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain every kernel launch on an NSys profile.

  • Design and implement exotic parallelism schemes to work with "interesting" hardware topologies.

  • Write custom GPU kernels to excel in specific regimes, such as cascade attention.

What we’re looking for
  • Strong understanding of LLM mechanics, like KV cache, mixture-of-experts, prefill vs. decode phases.

  • Interest in MLSys research—great ideas like speculative decoding and sparse attention come from research, that we need to follow closely.

  • Familiarity with modern, tile-based GPU programming, e.g. Triton, CUTLASS, ThunderKittens, etc. Or an interest in learning these!

Benefits

Meals are provided. Every employee receives a Studio Display.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Inference
Member of Technical Staff - Inference

Sail • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Inference Engineer, GPU Kernel Optimization
Senior Inference Engineer, GPU Kernel Optimization

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 288,000
Equity
Comprehensive benefits
Senior LLM Inference Engineer — Performance & GPU Optimization
Senior LLM Inference Engineer — Performance & GPU Optimization

Confidential • United States

On-site
USD 180,000 - 240,000
Machine Learning Engineer, LLM Inference Optimization in Sonoma
Machine Learning Engineer, LLM Inference Optimization in Sonoma

NLP PEOPLE • Sonoma (CA)

On-site
USD 120,000 - 160,000
Software Engineer – GPU Kernel
Software Engineer – GPU Kernel

FriendliAI • San Francisco (CA)

On-site
USD 120,000 - 150,000
Flexible working hours
Daily lunch and dinner
Health check-up support
+3
Member of Technical Staff, Kernels
Member of Technical Staff, Kernels

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
Machine Learning Engineer, LLM Inference Optimization
Machine Learning Engineer, LLM Inference Optimization

GMI Cloud • San Francisco (CA)

On-site
USD 180,000 - 240,000
MTS Inference: GPU Kernel & Performance Architect
MTS Inference: GPU Kernel & Performance Architect

Sail Research • San Francisco (CA)

On-site
USD 120,000 - 160,000
Meals provided
Studio Display for every employee
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Causal • San Francisco (CA)

On-site
USD 180,000 - 240,000