Inference Systems Performance Engineer

Adaption Labs

San Francisco, Northern (CA, KY)

Hybrid

USD 180,000 - 260,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Annual travel stipend
Lunch stipend
Well-Being benefits

Job summary

Adaption Labs is seeking a performance-focused ML systems engineer to own the cost and throughput of our inference stack. You will work with the serving fleet to optimize caching, batching, quantization, decoding, and kernel-level performance as workloads evolve.

You will collaborate with engineers, tune routing to external providers, and build profiling tools to reveal time and memory hot spots. Strong Python plus C++/Rust are highly valued.

Qualifications

  • 5+ years in ML systems, inference infra, or performance engineering with measurable improvements in cost or latency.
  • Deep understanding of model serving: prefill, decode, batching, memory bandwidth.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Strong Python skills and proficiency in C++, Rust, or another systems language.
  • Experience with GPU performance: CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.
  • Above all, we're looking for great teammates who are adaptable and bold.

Responsibilities

  • Improve throughput, cost, and tail latency via KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads based on real production traffic.
  • Tune routing between our infrastructure and external providers based on cost, capacity, and performance.
  • Work within serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
  • Build profiling and measurement systems that show where time, memory, and compute are being spent.

Skills

ML systems
Inference infra
Performance engineering
Python
C++/Rust
CUDA

Tools

vLLM
SGLang
TensorRT-LLM

Job description

Adaption Labs is seeking a performance-focused ML systems engineer to own the cost and throughput of our inference stack. You will work with the serving fleet to optimize caching, batching, quantization, decoding, and kernel-level performance as workloads evolve.

You will collaborate with engineers, tune routing to external providers, and build profiling tools to reveal time and memory hot spots. Strong Python plus C++/Rust are highly valued.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer: AI Serving at Scale
Inference Performance Engineer: AI Serving at Scale

adaption • United States

Hybrid
USD 180,000 - 240,000
Flexible in-person collaboration in BA
Adaption Passport travel stipend
Lunch stipend
+1
Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
Performance Engineer — AI Inference Systems
Performance Engineer — AI Inference Systems

Anthropic • San Francisco (CA)

Hybrid
USD 350,000 - 850,000
Visa sponsorship
Flexible hybrid work policy
Senior Inference Systems Engineer for High-Performance AI
Senior Inference Systems Engineer for High-Performance AI

Causal Labs • San Francisco (CA)

On-site
USD 150,000 - 210,000
Inference Performance Engineer
Inference Performance Engineer

adaption • United States

Hybrid
USD 180,000 - 240,000
Flexible in-person collaboration in BA
Adaption Passport travel stipend
Lunch stipend
+1
Inference Systems Engineer — High-Performance ML Runtime
Inference Systems Engineer — High-Performance ML Runtime

The Consensus • San Jose (CA)

On-site
USD 180,000 - 240,000
Medical/dental/vision benefits
Housing subsidy
Relocation support
+2
Inference Performance Engineer
Inference Performance Engineer

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
Senior Inference Systems Engineer — Low-Latency ML Serving
Senior Inference Systems Engineer — Low-Latency ML Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Inference Performance Engineer
Inference Performance Engineer

Adaption Labs • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 260,000
Annual travel stipend
Lunch stipend
Well-Being benefits
Inference Performance Engineer: Benchmark & Optimize
Inference Performance Engineer: Benchmark & Optimize

Coreweave • Bellevue (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+1