Inference Performance Engineer: Optimize ML Serving & Latency (Flexible Work)

adaption

Dublin

On-site

EUR 120,000 - 180,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible work
Adaption Passport
Lunch stipend
Well-being benefits

Job summary

Adaption is seeking an senior ML systems engineer to own the cost and performance of the inference stack. You will optimize caching, batching, and decoding while coordinating with the serving fleet to improve throughput and latency without compromising model quality.

You will work with engines like vLLM, SGLang, and TensorRT-LLM, implement profiling tools, and help route workloads to balance cost and performance in a global deployment.

Qualifications

  • 5+ years in ML systems, inference infrastructure, or performance engineering with measurable improvements in cost or latency.
  • Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Strong Python skills and proficiency in C++, Rust, or another systems language.
  • Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.

Responsibilities

  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads based on real production traffic.
  • Tune routing between our infrastructure and external providers based on cost, capacity, and performance.
  • Work within serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
  • Build profiling and measurement systems that show where time, memory, and compute are being spent.

Skills

Python
C++
Rust
Performance engineering

Tools

vLLM
SGLang
TensorRT-LLM
CUDA
NCCL

Job description

Adaption is seeking an senior ML systems engineer to own the cost and performance of the inference stack. You will optimize caching, batching, and decoding while coordinating with the serving fleet to improve throughput and latency without compromising model quality.

You will work with engines like vLLM, SGLang, and TensorRT-LLM, implement profiling tools, and help route workloads to balance cost and performance in a global deployment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer
Inference Performance Engineer

adaption • Dublin

On-site
EUR 120,000 - 180,000
Flexible work
Adaption Passport
Lunch stipend
+1
AI Inference Engineer — High-Performance, Low-Latency ML
AI Inference Engineer — High-Performance, Low-Latency ML

F5 • Dublin

On-site
EUR 90,000 - 150,000
Senior AI Inference Engineer - High-Throughput LLM Serving
Senior AI Inference Engineer - High-Throughput LLM Serving

Confidential • Ireland

On-site
EUR 120,000 - 180,000
Senior ML Inference Systems Engineer (Hybrid, On-Prem GPU)
Senior ML Inference Systems Engineer (Hybrid, On-Prem GPU)

Uniting Holding • Dublin

Hybrid
EUR 80,000 - 120,000
25 days paid annual leave
Free inference tokens
Remote work flexibility
Senior ML Systems Engineer (Inference)
Senior ML Systems Engineer (Inference)

Uniting Holding • Dublin

Hybrid
EUR 80,000 - 120,000
25 days paid annual leave
Free inference tokens
Remote work flexibility
Senior LLM Inference & GPU Performance Engineer
Senior LLM Inference & GPU Performance Engineer

Confidential • Ireland

On-site
EUR 70,000 - 90,000
Senior Engineer - AI, Inference
Senior Engineer - AI, Inference

Confidential • Ireland

On-site
EUR 120,000 - 180,000
Staff AI Model Optimization Architect — Scalable Inference
Staff AI Model Optimization Architect — Scalable Inference

Qualcomm • Cork

On-site
EUR 150,000 - 190,000
Salary and stock bonus
Relocation assistance
Education assistance
+4
LLM Inference Engineer — Low-Latency, High-Throughput
LLM Inference Engineer — Low-Latency, High-Throughput

F5 Networks, Inc.  • Dublin

On-site
EUR 70,000 - 90,000
Flexible work conditions
Equal employment opportunities
Senior AI Model Optimization Architect for Inference
Senior AI Model Optimization Architect for Inference

Qualcomm • Ireland

On-site
EUR 120,000 - 180,000
Salary, stock and performance related—
Relocation and immigration support
Education Assistance
+1