ML Inference Performance Engineer — Optimize Cost & Latency

Adaption Labs

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Flexible work
Adaption Passport
Lunch stipend
Well-Being benefits

Job summary

Adaption Labs in San Francisco is looking for a senior ML systems engineer to own the cost and performance of the inference stack. You’ll optimize caching, batching, quantization, and kernel-level performance to improve throughput and reduce latency without sacrificing model quality.

You’ll collaborate with the serving fleet engineers and tune routing, profiling, and measurement systems across production workloads. A strong background in Python/C++/Rust and GPU optimization is required.

Qualifications

  • 5+ years in ML systems, inference infrastructure, or performance engineering.
  • Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Strong Python skills and proficiency in C++, Rust, or another systems language.
  • Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.

Responsibilities

  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads based on real production traffic.
  • Tune routing between our infrastructure and external providers based on cost, capacity, and performance.
  • Work within serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
  • Build profiling and measurement systems that show where time, memory, and compute are being spent.

Skills

ML systems
Model serving
Python
C++ / Rust
GPU performance

Tools

vLLM
SGLang
TensorRT-LLM

Job description

Adaption Labs in San Francisco is looking for a senior ML systems engineer to own the cost and performance of the inference stack. You’ll optimize caching, batching, quantization, and kernel-level performance to improve throughput and reduce latency without sacrificing model quality.

You’ll collaborate with the serving fleet engineers and tune routing, profiling, and measurement systems across production workloads. A strong background in Python/C++/Rust and GPU optimization is required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,000 - 263,000
Health insurance
401(k) plan
Parental leave
+2
Low-Latency ML Inference Engineer
Low-Latency ML Inference Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
ML Inference Systems Engineer
ML Inference Systems Engineer

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
INFERENCE OPTIMIZATION ENGINEER
INFERENCE OPTIMIZATION ENGINEER

Up Top • United States

Hybrid
USD 180,000 - 320,000
ML Inference Performance Visibility Engineer
ML Inference Performance Visibility Engineer

Etched • San Jose (CA)

On-site
USD 150,000 - 210,000
Housing subsidy
Relocation support
Medical benefits
+1
Performance Engineer, Inference Engine - High-Performance AI
Performance Engineer, Inference Engine - High-Performance AI

EngineersOfAI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
High-Performance ML Inference Engineer
High-Performance ML Inference Engineer

Reactor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive SF salary
Early equity
Visa sponsorship
+2