Inference Performance Engineer: Optimize Model Serving

Adaption

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
Flexible in-person collaboration

Job summary

Adaption in San Francisco Bay Area seeks a senior ML systems engineer to own cost and performance of the inference stack. You will optimize caching, batching, quantization, and decoding, while collaborating with the serving fleet to deliver scalable, low-latency model serving.

Required are 5+ years in ML systems with deep knowledge of model serving, GPU performance, and proficiency in Python and C++/Rust. You’ll work with engines like vLLM and TensorRT-LLM and contribute to profiling and

Qualifications

  • 5+ years in ML systems, inference infrastructure, or performance engineering with measurable improvements in cost or latency.
  • Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Strong Python skills and proficiency in C++, Rust, or another systems language.
  • Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.

Responsibilities

  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads based on real production traffic.
  • Tune routing between our infrastructure and external providers based on cost, capacity, and performance.
  • Work within serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
  • Build profiling and measurement systems that show where time, memory, and compute are being spent.

Skills

ML systems
Performance engineering
Python
C++/Rust
GPU performance

Tools

vLLM
SGLang
TensorRT-LLM
CUDA
NCCL

Job description

Adaption in San Francisco Bay Area seeks a senior ML systems engineer to own cost and performance of the inference stack. You will optimize caching, batching, quantization, and decoding, while collaborating with the serving fleet to deliver scalable, low-latency model serving.

Required are 5+ years in ML systems with deep knowledge of model serving, GPU performance, and proficiency in Python and C++/Rust. You’ll work with engines like vLLM and TensorRT-LLM and contribute to profiling and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Systems Performance Engineer
Inference Systems Performance Engineer

Adaption Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Annual travel stipend
Lunch stipend
Well-Being benefits
Inference Performance Engineer: AI Serving at Scale
Inference Performance Engineer: AI Serving at Scale

adaption • United States

Hybrid
USD 180,000 - 240,000
Flexible in-person collaboration in BA
Adaption Passport travel stipend
Lunch stipend
+1
Senior Inference Systems Engineer — Low-Latency ML Serving
Senior Inference Systems Engineer — Low-Latency ML Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,000 - 263,000
Health insurance
401(k) plan
Parental leave
+2
ML Inference Engineer San Francisco · Engineering · Full Time →
ML Inference Engineer San Francisco · Engineering · Full Time →

Reactor • San Francisco (CA)

On-site
USD 120,000 - 160,000
Visa sponsorship
Relocation support
Generous health, dental, and vision coverage
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Inference Performance Engineer
Inference Performance Engineer

adaption • United States

Hybrid
USD 180,000 - 240,000
Flexible in-person collaboration in BA
Adaption Passport travel stipend
Lunch stipend
+1