Inference Engineer: High-Throughput ML Serving

Adaption Labs, Inc.

San Francisco (CA)

Hybrid

USD 180,000 - 230,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Flexible work: In-person collaboration
Adaption Passport
Lunch Stipend
Well-Being

Job summary

Adaption Labs, Inc. is seeking an ML systems engineer to own the cost and performance of the inference stack. You will optimize caching, batching, quantization, decoding, and kernel-level tuning to improve throughput and latency while preserving model quality.

You will collaborate with the serving fleet engineers and tackle real production workloads, focusing on cost-efficiency, tail latency, and reliable delivery across changing workloads and hardware. Bay Area presence required.

Qualifications

  • 5+ years in ML systems, inference infrastructure, or performance engineering with measurable improvements in cost or latency.
  • Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Strong Python skills and proficiency in C++, Rust, or another systems language.
  • Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.

Responsibilities

  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads based on real production traffic.
  • Tune routing between our infrastructure and external providers based on cost, capacity, and performance.
  • Work within serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
  • Build profiling and measurement systems that show where time, memory, and compute are being spent.

Skills

ML systems
Python
Performance engineering

Tools

vLLM
SGLang
TensorRT-LLM
CUDA

Job description

Adaption Labs, Inc. is seeking an ML systems engineer to own the cost and performance of the inference stack. You will optimize caching, batching, quantization, decoding, and kernel-level tuning to improve throughput and latency while preserving model quality.

You will collaborate with the serving fleet engineers and tackle real production workloads, focusing on cost-efficiency, tail latency, and reliable delivery across changing workloads and hardware. Bay Area presence required.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Inference Performance Engineer — Optimize Cost & Latency
ML Inference Performance Engineer — Optimize Cost & Latency

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
Inference Systems Engineer: Optimize AI Serving & Latency
Inference Systems Engineer: Optimize AI Serving & Latency

Emploive • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Flexible work
Adaption Passport
Lunch Stipend
+1
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
Staff ML Engineer - Efficient Production Inference
Staff ML Engineer - Efficient Production Inference

ATBF Labs • San Francisco (CA)

Hybrid
USD 215,000 - 285,000
Equity
Health benefits
Inference Engineer Adaption · San Francisco, CA Full-time · Hybrid — 41 minutes ago
Inference Engineer Adaption · San Francisco, CA Full-time · Hybrid — 41 minutes ago

Emploive • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Flexible work
Adaption Passport
Lunch Stipend
+1
LLM Inference Performance Engineer: Speed & Efficiency
LLM Inference Performance Engineer: Speed & Efficiency

Baseten • United States

Remote
USD 150,000 - 210,000
ML Inference Systems Engineer — High-Performance Serving
ML Inference Systems Engineer — High-Performance Serving

Annapurna Labs (U.S.) Inc. • Seattle (WA)

On-site
USD 144,000 - 194,000
RSUs
401(k) matching
Paid time off
Senior ML Inference Engineer – High-Performance Serving
Senior ML Inference Engineer – High-Performance Serving

Amazon Inc. • Seattle (WA)

On-site
USD 168,000 - 227,000
Inference Engineer
Inference Engineer

Adaption Labs, Inc. • San Francisco (CA)

Hybrid
USD 180,000 - 230,000
Flexible work: In-person collaboration
Adaption Passport
Lunch Stipend
+1
Senior ML Inference Systems Engineer
Senior ML Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 180,000 - 240,000