Inference Systems Engineer: Optimize AI Serving & Latency

adaption

San Francisco (CA)

Hybrid

USD 180,000 - 240,000

Full time

22 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Flexible work in Bay Area
Adaption Passport travel stipend
Lunch stipend
Medical benefits & PTO

Job summary

adaption is seeking an experienced engineer to own the cost and performance of our inference stack in the San Francisco Bay Area. You will influence throughput, latency, and reliability as workloads and hardware evolve.

You'll collaborate with the serving fleet engineers and own core levers like caching, batching, quantization, and decoding, ensuring efficient model serving without sacrificing quality.

Qualifications

  • 5+ years in ML systems, inference infrastructure, or performance engineering.
  • Experience with model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.

Responsibilities

  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads based on real production traffic.
  • Tune routing between our infrastructure and external providers based on cost, capacity, and performance.
  • Work within serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
  • Build profiling and measurement systems that show where time, memory, and compute are being spent.

Skills

Python
C++
Rust
ML systems
Performance engineering
GPU performance

Tools

vLLM
SGLang
TensorRT-LLM

Job description

adaption is seeking an experienced engineer to own the cost and performance of our inference stack in the San Francisco Bay Area. You will influence throughput, latency, and reliability as workloads and hardware evolve.

You'll collaborate with the serving fleet engineers and own core levers like caching, batching, quantization, and decoding, ensuring efficient model serving without sacrificing quality.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference Systems Engineer: Optimize AI Serving & Latency
Inference Systems Engineer: Optimize AI Serving & Latency

Emploive • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Flexible work
Adaption Passport
Lunch Stipend
+1
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
Inference Performance Engineer: Latency & Cost Optimization
Inference Performance Engineer: Latency & Cost Optimization

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 555,000
Inference Engineer Adaption · San Francisco, CA Full-time · Hybrid — 41 minutes ago
Inference Engineer Adaption · San Francisco, CA Full-time · Hybrid — 41 minutes ago

Emploive • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Flexible work
Adaption Passport
Lunch Stipend
+1
Senior Inference Systems Engineer for High-Performance AI
Senior Inference Systems Engineer for High-Performance AI

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Inference Systems Engineer — High-Performance AI Serving
Inference Systems Engineer — High-Performance AI Serving

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Inference Systems Engineer: Scalable AI Infrastructure
Inference Systems Engineer: Scalable AI Infrastructure

Candidate • San Francisco (CA)

On-site
USD 180,000 - 240,000
Flexible PTO
Medical, dental, vision benefits
Life insurance and disability benefits
+4
Inference Performance Engineer for Visual AI
Inference Performance Engineer for Visual AI

AI Chopping Block, Inc. • San Francisco (CA)

On-site
USD 150,000 - 230,000
Equity
401k
Healthcare
+1
Senior Systems Engineer, AI Inference Platform
Senior Systems Engineer, AI Inference Platform

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 260,000
AI Inference Engineer - Scalable, Low-Latency Systems
AI Inference Engineer - Scalable, Low-Latency Systems

SPACE EXPLORATION TECHNOLOGIES CORP • Palo Alto (CA), Northern (KY)

Hybrid
USD 135,000 - 210,000
401(k)
Medical, vision and dental coverage
Paid parental leave
+4