Senior AI Inference Engineer - High-Throughput LLM Serving

Confidential

Ireland

On-site

EUR 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Confidential in Ireland is seeking a hands-on ML engineering lead to optimise LLM inference pipelines and own end-to-end model serving in production. You will deploy multi-GPU inference, manage low-latency pathways, and build OpenAI-compatible APIs across cloud and on-prem environments.

You should have deep experience with vLLM/TensorRT-LLM/TGI, strong GPU performance skills, and solid Python/GoLang software engineering.

Qualifications

  • Hands-on production experience serving LLMs at scale with measurable throughput/latency improvements.
  • Deep familiarity with a modern inference/serving framework (vLLM, TensorRT-LLM, TGI, or similar).
  • Strong grip on GPU performance: memory management and model/tensor parallelism.
  • Solid software engineering in Python or GoLang, plus Docker + Kubernetes for production deployment.
  • A benchmarking mindset—measure, compare, and defend trade-offs.

Responsibilities

  • Optimise LLM inference pipelines (multi-GPU inference, prefix caching, memory-efficient serving).
  • Own end-to-end model serving in production (deployment, low-latency inference, multi-GPU parallelism).
  • Build and maintain OpenAI-compatible serving APIs (/v1/chat, /v1/responses).
  • Tune and operate a modern serving stack (vLLM) with batching and cache management.
  • Maximise GPU usage across architectures; profile and remove bottlenecks.
  • Instrument the serving layer with logging, telemetry, and metrics for observability and autoscaling.
  • Ship on Kubernetes: Docker, Helm, CI/CD, staged rollouts.
  • Benchmark rigorously against standards and build tooling for performance characterization.

Skills

Production experience
LLM inference
GPU performance
Python
Go
Docker
Kubernetes
Benchmarking

Tools

vLLM
TensorRT-LLM
TGI

Job description

Confidential in Ireland is seeking a hands-on ML engineering lead to optimise LLM inference pipelines and own end-to-end model serving in production. You will deploy multi-GPU inference, manage low-latency pathways, and build OpenAI-compatible APIs across cloud and on-prem environments.

You should have deep experience with vLLM/TensorRT-LLM/TGI, strong GPU performance skills, and solid Python/GoLang software engineering.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior LLM Inference & GPU Performance Engineer
Senior LLM Inference & GPU Performance Engineer

Confidential • Ireland

On-site
EUR 70,000 - 90,000
LLM Inference Engineer — Low-Latency, High-Throughput
LLM Inference Engineer — Low-Latency, High-Throughput

F5 Networks, Inc.  • Dublin

On-site
EUR 70,000 - 90,000
Flexible work conditions
Equal employment opportunities
Senior AI SRE: LLM Infra & GPU-Powered Ops
Senior AI SRE: LLM Infra & GPU-Powered Ops

Confidential • Dublin

On-site
EUR 120,000 - 180,000
LLM Inference Engineer — Cloud AI Platform
LLM Inference Engineer — Cloud AI Platform

Qualcomm • Ireland

On-site
EUR 110,000 - 170,000
Salary and bonus
Parental leave
Employee stock purchase
+4
LLM Serving Engineer, Cloud AI Platform
LLM Serving Engineer, Cloud AI Platform

Qualcomm • Cork

Hybrid
EUR 110,000 - 160,000
Salary and stock options
Performance bonus
Maternity/Paternity Leave
+8
AI Inference Engineer — High-Performance, Low-Latency ML
AI Inference Engineer — High-Performance, Low-Latency ML

F5 • Dublin

On-site
EUR 90,000 - 150,000
Inference Performance Engineer: Optimize ML Serving & Latency (Flexible Work)
Inference Performance Engineer: Optimize ML Serving & Latency (Flexible Work)

adaption • Dublin

On-site
EUR 120,000 - 180,000
Flexible work
Adaption Passport
Lunch stipend
+1
Senior ML Systems Engineer (Inference)
Senior ML Systems Engineer (Inference)

Uniting Holding • Dublin

Hybrid
EUR 80,000 - 120,000
25 days paid annual leave
Free inference tokens
Remote work flexibility
GenAI/LLM Engineer – On-Prem GPU, Remote Ireland
GenAI/LLM Engineer – On-Prem GPU, Remote Ireland

Ergo IT Recruitment Services • Dublin

On-site
EUR 90,000 - 120,000
Senior Engineer - AI, Inference
Senior Engineer - AI, Inference

Confidential • Ireland

On-site
EUR 120,000 - 180,000