Inference Performance Engineer: Latency & Cost Optimizer

adaption

Singapore

Hybrid

SGD 120,000 - 160,000

Full time

48 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible work
Adaption Passport
Lunch Stipend
Well-Being

Job summary

adaption is seeking an experienced ML systems engineer to own the cost and performance of our inference stack, ensuring efficient model serving as workloads and hardware evolve. You will collaborate with the serving fleet to optimize caching, batching, quantization, decoding, and kernel-level tuning for higher throughput and lower latency without sacrificing reliability or model quality.

You will drive improvements in long-context prefill and decode workloads, optimize routing to external

Qualifications

  • 5+ years in ML systems, inference infrastructure, or performance engineering with measurable improvements in cost or latency.
  • Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Strong Python skills and proficiency in C++, Rust, or another systems language.
  • Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.

Responsibilities

  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads based on real production traffic.
  • Tune routing between our infrastructure and external providers based on cost, capacity, and performance.
  • Work within serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
  • Build profiling and measurement systems that show where time, memory, and compute are being spent.

Skills

ML systems
Inference infrastructure
Performance engineering
Python
C++

Tools

vLLM
SGLang
TensorRT-LLM

Job description

adaption is seeking an experienced ML systems engineer to own the cost and performance of our inference stack, ensuring efficient model serving as workloads and hardware evolve. You will collaborate with the serving fleet to optimize caching, batching, quantization, decoding, and kernel-level tuning for higher throughput and lower latency without sacrificing reliability or model quality.

You will drive improvements in long-context prefill and decode workloads, optimize routing to external

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer
Inference Performance Engineer

adaption • Singapore

Hybrid
SGD 120,000 - 160,000
Flexible work
Adaption Passport
Lunch Stipend
+1
Senior LLM Inference Performance Engineer
Senior LLM Inference Performance Engineer

Bitdeer Group • Singapore

On-site
SGD 150,000 - 210,000
Senior LLM Inference Engineer Performance & GPU Optimization
Senior LLM Inference Engineer Performance & GPU Optimization

Confidential • Singapore

On-site
SGD 90,000 - 130,000
Machine Learning Systems Senior Engineer
Machine Learning Systems Senior Engineer

Defence Science and Technology Agency • Singapore

On-site
SGD 120,000 - 180,000
Production-Scale LLM Inference Performance Engineer
Production-Scale LLM Inference Performance Engineer

Confidential • Singapore

On-site
SGD 90,000 - 130,000
Senior ML Systems Engineer - Inference Platforms
Senior ML Systems Engineer - Inference Platforms

Defence Science and Technology Agency • Singapore

On-site
SGD 120,000 - 180,000
Senior LLM Inference Performance Engineer
Senior LLM Inference Performance Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
Welfare benefits
Training and mentoring
High-Performance Backend Inference Engineer
High-Performance Backend Inference Engineer

Bytedance • Singapore

On-site
SGD 180,000 - 280,000
High-Performance ML Backend Engineer
High-Performance ML Backend Engineer

United States Digital Space LLC • Singapore

On-site
SGD 70,000 - 120,000
Senior Inference Runtime Engineer
Senior Inference Runtime Engineer

Bitdeer Group • Singapore

On-site
SGD 150,000 - 210,000