Senior Engineer - AI, Inference

Confidential

Ireland

On-site

EUR 120,000 - 180,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Confidential in Ireland is seeking a hands-on ML engineering lead to optimise LLM inference pipelines and own end-to-end model serving in production. You will deploy multi-GPU inference, manage low-latency pathways, and build OpenAI-compatible APIs across cloud and on-prem environments.

You should have deep experience with vLLM/TensorRT-LLM/TGI, strong GPU performance skills, and solid Python/GoLang software engineering.

Qualifications

  • Hands-on production experience serving LLMs at scale with measurable throughput/latency improvements.
  • Deep familiarity with a modern inference/serving framework (vLLM, TensorRT-LLM, TGI, or similar).
  • Strong grip on GPU performance: memory management and model/tensor parallelism.
  • Solid software engineering in Python or GoLang, plus Docker + Kubernetes for production deployment.
  • A benchmarking mindset—measure, compare, and defend trade-offs.

Responsibilities

  • Optimise LLM inference pipelines (multi-GPU inference, prefix caching, memory-efficient serving).
  • Own end-to-end model serving in production (deployment, low-latency inference, multi-GPU parallelism).
  • Build and maintain OpenAI-compatible serving APIs (/v1/chat, /v1/responses).
  • Tune and operate a modern serving stack (vLLM) with batching and cache management.
  • Maximise GPU usage across architectures; profile and remove bottlenecks.
  • Instrument the serving layer with logging, telemetry, and metrics for observability and autoscaling.
  • Ship on Kubernetes: Docker, Helm, CI/CD, staged rollouts.
  • Benchmark rigorously against standards and build tooling for performance characterization.

Skills

Production experience
LLM inference
GPU performance
Python
Go
Docker
Kubernetes
Benchmarking

Tools

vLLM
TensorRT-LLM
TGI

Job description

We build the inference layer that powers F5's AI security products: the systems that run large language models fast, cheaply, and reliably at production scale so we can inspect, secure, and govern enterprise giveGenAI traffic in real time.

You’ll own how models are served — squeezing maximum throughput out of every GPU, cutting tail latency, and keeping the serving stack observable and self-scaling under real customer load. If you think in tokens-per-second, prefill latency, and GPU memory budgets, this is your seat.

What you'll do
  • Optimise LLM inference pipelines (LLaMA, GLM, GPT-OSS and similar model classes) for throughput and latency — multi-GPU inference, prefix caching, and memory-efficient serving — targeting order-of-magnitude gains in requests-per-second (RPS).
  • Own end-to-end model serving in production: deployment, low-latency inference, multi gpu parallelism, and high-throughput serving across cloud and on-prem environments.
  • Build and maintain OpenAI-compatible serving APIs (e.g. /v1/chat, /v1/responses) that support reliable tool calling across multi-step agentic workflows.
  • Tune and operate a modern serving stack (vLLM or equivalent) — continuous batching, KV-cache management — balancing throughput, latency, generation quality, and memory footprint.
  • Maximise GPU across chip architectures like Ada Lovelace, Hopper, Blackwell, etc; profile, benchmark, and eliminate bottlenecks.
  • Instrument the serving layer with logging, telemetry, and metrics (prefill latency, tokens-per-batch capacity, preemption count, queue depth) to drive observability and metric-based autoscaling.
  • Ship on Kubernetes: containerised deployments (Docker, Helm), CI/CD, and low-risk, version-controlled rollouts across staging and production.
  • Benchmark rigorously against recognised standards and build tooling that automates performance characterisation.
What you'll bring (must-have)
  • Hands-on production experience serving LLMs at scale, with measurable throughput/latency wins you can walk through end to end.
  • Deep familiarity with a modern inference/serving framework (vLLM, TensorRT-LLM, TGI, or similar), including batching and speculative decoding.
  • Strong grip on GPU performance: memory management, model/tensor parallelism, and hardware-aware optimisation.
  • Solid software engineering in Python OR GoLang, plus Docker + Kubernetes for production deployment.
  • A benchmarking mindset — you measure rather than guess, and you can defend the trade-offs.
Nice to have
  • Building OpenAI-compatible endpoints and agentic / tool-calling serving paths.
  • Distributed training exposure (large-model fine-tuning / pretraining on multi-node GPU clusters).
  • MLPerf submission experience or embedded / ARM inference optimisation.
  • Release-management / CI-CD ownership across production and staging.
  • Interest in AI security — adversarial robustness, model scanning, or securing GenAI in production.
Tech you'll work with
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer (Inference)
Senior ML Systems Engineer (Inference)

Uniting Holding • Dublin

Hybrid
EUR 80,000 - 120,000
25 days paid annual leave
Free inference tokens
Remote work flexibility
AI Inference Engineer
AI Inference Engineer

F5 • Dublin

On-site
EUR 90,000 - 150,000
Senior Site Reliability Engineer, AI Inference
Senior Site Reliability Engineer, AI Inference

F5 Networks, Inc.  • Dublin

Hybrid
EUR 120,000 - 180,000
SIte Reliability Engineer III, AI Inference
SIte Reliability Engineer III, AI Inference

F5 Networks, Inc.  • Dublin

Hybrid
EUR 90,000 - 130,000
AI Inference Engineer — High-Performance, Low-Latency ML
AI Inference Engineer — High-Performance, Low-Latency ML

F5 • Dublin

On-site
EUR 90,000 - 150,000
Senior AI Inference Engineer - High-Throughput LLM Serving
Senior AI Inference Engineer - High-Throughput LLM Serving

Confidential • Ireland

On-site
EUR 120,000 - 180,000
Senior AI Inference SRE: Low-Latency, Scalable Serving
Senior AI Inference SRE: Low-Latency, Scalable Serving

F5 Networks, Inc.  • Dublin

Hybrid
EUR 90,000 - 130,000
Hybrid/Remote Senior AI Inference Engineer
Hybrid/Remote Senior AI Inference Engineer

F5 Networks, Inc.  • Dublin

Hybrid
EUR 120,000 - 180,000
Staff Software Engineer, Inference
Staff Software Engineer, Inference

Anthropic • Dublin

Hybrid
EUR 80,000 - 110,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
AI Inference Engineer
AI Inference Engineer

F5 Networks, Inc.  • Dublin

On-site
EUR 70,000 - 90,000
Flexible work conditions
Equal employment opportunities