Inference Performance Engineer | Optimize AI Serving

adaption

Dublin

Hybrid

EUR 120,000 - 180,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Flexible work
Adaption Passport travel stipend
Lunch stipend
Well-Being benefits

Job summary

Adaption is seeking an experienced ML systems engineer who will own the cost and performance of our inference stack. You will work with the serving fleet team to optimize caching, batching, quantization, and kernel-level tuning to maximize throughput and minimize tail latency without sacrificing model quality.

You will tune routing between internal infrastructure and external providers and build profiling tools to illuminate time and memory usage.

Qualifications

  • 5+ years in ML systems, inference infrastructure, or performance engineering.
  • Deep understanding of model serving, including prefill and decode, memory bandwidth, batching, and concurrency.
  • Production experience with serving engines such as vLLM, SGLang, or TensorRT-LLM.
  • Strong Python skills and proficiency in C++, Rust, or another systems language.
  • Experience with GPU performance, including CUDA, NCCL, mixed precision, memory layout, kernels, or quantization.

Responsibilities

  • Improve throughput, cost, and tail latency through KV-cache management, continuous batching, speculative decoding, and quantization.
  • Optimize long-context prefill and decode workloads based on real production traffic.
  • Tune routing between our infrastructure and external providers based on cost, capacity, and performance.
  • Work within serving engines such as vLLM, SGLang, and TensorRT-LLM, going below the framework when needed.
  • Build profiling and measurement systems that show where time, memory, and compute are being spent.

Skills

ML systems
Performance engineering
Python
C++
Rust
GPU performance

Tools

vLLM
SGLang
TensorRT-LLM

Job description

Adaption is seeking an experienced ML systems engineer who will own the cost and performance of our inference stack. You will work with the serving fleet team to optimize caching, batching, quantization, and kernel-level tuning to maximize throughput and minimize tail latency without sacrificing model quality.

You will tune routing between internal infrastructure and external providers and build profiling tools to illuminate time and memory usage.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Performance Engineer
Inference Performance Engineer

adaption • Dublin

Hybrid
EUR 120,000 - 180,000
Flexible work
Adaption Passport travel stipend
Lunch stipend
+1
AI Inference Performance Engineer
AI Inference Performance Engineer

Nutanix • Cork

On-site
EUR 90,000 - 150,000
Stock options
Performance bonus
Relocation support
+1
Senior AI Inference Engineer - High-Throughput LLM Serving
Senior AI Inference Engineer - High-Throughput LLM Serving

Confidential • Ireland

On-site
EUR 120,000 - 180,000
AI Infrastructure Engineer — Scalable ML Systems
AI Infrastructure Engineer — Scalable ML Systems

Jobtailor • Leinster

On-site
EUR 90,000 - 120,000
AI Inference Performance Engineer
AI Inference Performance Engineer

Qualcomm • Cork

Hybrid
EUR 90,000 - 130,000
Salary review and performance bonus
Relocation support
Education Assistance
+2
Staff Engineer, Inference at Scale
Staff Engineer, Inference at Scale

Anthropic • Dublin

Hybrid
EUR 120,000 - 180,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
+1
Senior ML Inference Systems Engineer (Hybrid, On-Prem GPU)
Senior ML Inference Systems Engineer (Hybrid, On-Prem GPU)

Uniting Holding • Dublin

Hybrid
EUR 80,000 - 120,000
25 days paid annual leave
Free inference tokens
Remote work flexibility
Senior Engineer - AI, Inference
Senior Engineer - AI, Inference

Confidential • Ireland

On-site
EUR 120,000 - 180,000
Senior AI Model Optimization Architect for Inference
Senior AI Model Optimization Architect for Inference

Qualcomm • Ireland

On-site
EUR 120,000 - 180,000
Salary, stock and performance related—
Relocation and immigration support
Education Assistance
+1
Cloud AI Performance Engineer
Cloud AI Performance Engineer

Qualcomm • Ireland

On-site
EUR 90,000 - 150,000
Stock bonus
Employee stock purchase scheme
Pension matching scheme
+4