Inference Performance Engineer: AI Serving at Scale

adaption

United States

Hybrid

USD 180,000 - 240,000

Full time

7 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible in-person collaboration in BA
Adaption Passport travel stipend
Lunch stipend
Well-being: medical benefits and PTO

Job summary

Adaption is seeking an experienced engineer to own the cost and performance of our inference stack in a production environment. You will influence batching, caching, quantization, and kernel-level optimization while coordinating with the serving fleet to meet throughput and latency targets.

You will optimize prefill/decode workloads, route performance between internal and external providers, and build profiling tools.

Qualifications

  • 5+ years in ML systems, inference infrastructure, or performance engineering with measurable improvements in cost or latency.
  • Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Strong Python skills and proficiency in C++, Rust, or another systems language.
  • Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.

Responsibilities

  • Own throughput, cost, and tail latency improvements through KV-cache management, batching, quantization, and decoding.
  • Optimize long-context prefill and decode workloads based on real production traffic.
  • Tune routing between our infrastructure and external providers by cost, capacity, and performance.
  • Work with serving engines such as vLLM, SGLang, and TensorRT-LLM, and go below the framework when needed.
  • Build profiling and measurement systems to show where time, memory, and compute are spent.

Skills

ML systems
Inference infra
Performance engineering
Python
C++
Rust
GPU performance
CUDA
NCCL
Quantization

Tools

vLLM
SGLang
TensorRT-LLM

Job description

Adaption is seeking an experienced engineer to own the cost and performance of our inference stack in a production environment. You will influence batching, caching, quantization, and kernel-level optimization while coordinating with the serving fleet to meet throughput and latency targets.

You will optimize prefill/decode workloads, route performance between internal and external providers, and build profiling tools.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Systems Performance Engineer
Inference Systems Performance Engineer

Adaption Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Annual travel stipend
Lunch stipend
Well-Being benefits
Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
Inference Performance Engineer: Benchmark & Optimize
Inference Performance Engineer: Benchmark & Optimize

Coreweave • Bellevue (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+1
Senior AI Infra Engineer: High-Performance Inference
Senior AI Infra Engineer: High-Performance Inference

Ddn • Sacramento (CA)

On-site
USD 140,000 - 200,000
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
Inference Performance Engineer - Benchmark & Optimize
Inference Performance Engineer - Benchmark & Optimize

CoreWeave • Sunnyvale (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+2
Senior AI Inference Performance Engineer — Scale GPUs
Senior AI Inference Performance Engineer — Scale GPUs

NVIDIA • California (MO)

On-site
USD 124,000 - 196,000
Equity eligibility
Applied AI Inference Performance Engineer
Applied AI Inference Performance Engineer

CoreWeave • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical benefits
401(k) with match
Flexible PTO
+5
Inference Performance Engineer
Inference Performance Engineer

adaption • United States

Hybrid
USD 180,000 - 240,000
Flexible in-person collaboration in BA
Adaption Passport travel stipend
Lunch stipend
+1
Inference Performance Engineer
Inference Performance Engineer

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1