Inference Performance Engineer — Scale AI Serving

Adaption Labs

United States

Hybrid

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible work
Adaption Passport travel stipend
Lunch stipend
Well-being benefits

Job summary

Adaption Labs seeks an engineer to own the cost and performance of our inference stack. You will shape how we serve models as workloads shift, focusing on throughput, latency, and reliability.

You will collaborate with the serving fleet engineers, tuning caching, batching, quantization, decoding, and kernel-level optimizations to maximize efficiency without compromising model quality.

Qualifications

  • 5+ years in ML systems, inference infrastructure, or performance engineering with measurable improvements in cost or latency.
  • Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Strong Python skills and proficiency in C++, Rust, or another systems language.
  • Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.

Responsibilities

  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads based on real production traffic.
  • Tune routing between our infrastructure and external providers based on cost, capacity, and performance.
  • Work within serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
  • Build profiling and measurement systems that show where time, memory, and compute are being spent.

Skills

ML systems
Performance engineering
Serving engines
Python/C++/Rust
GPU performance

Job description

Adaption Labs seeks an engineer to own the cost and performance of our inference stack. You will shape how we serve models as workloads shift, focusing on throughput, latency, and reliability.

You will collaborate with the serving fleet engineers, tuning caching, batching, quantization, decoding, and kernel-level optimizations to maximize efficiency without compromising model quality.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Systems Performance Engineer
Inference Systems Performance Engineer

Adaption Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Annual travel stipend
Lunch stipend
Well-Being benefits
Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
Distributed Inference Performance Engineer
Distributed Inference Performance Engineer

OpenAI • California (MO)

On-site
USD 150,000 - 190,000
Inference Performance Engineer: Latency & Cost Optimization
Inference Performance Engineer: Latency & Cost Optimization

OpenAI • San Francisco (CA)

On-site
USD 295,000 - 555,000
Inference Performance Engineer: Benchmark & Optimize
Inference Performance Engineer: Benchmark & Optimize

Coreweave • Bellevue (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+1
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
Staff Engineer – High-Performance Model Inference
Staff Engineer – High-Performance Model Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Inference Performance Engineer - Benchmark & Optimize
Inference Performance Engineer - Benchmark & Optimize

CoreWeave • Sunnyvale (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+2
Inference Infra Engineer: Scale Low-Latency AI Serving
Inference Infra Engineer: Scale Low-Latency AI Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Inference Performance Engineer
Inference Performance Engineer

Adaption Labs • United States

Hybrid
USD 140,000 - 210,000
Flexible work
Adaption Passport travel stipend
Lunch stipend
+1