Senior Inference Engineer — GPU-Disaggregated Serving

TensorX

Dublin

Hybrid

EUR 120,000 - 180,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Hybrid working
25 days paid annual leave
Free inference tokens!

Job summary

TensorX, a Dublin-based sovereign AI infrastructure platform, seeks a Senior Inference Engineer to own the disaggregated serving layer across multiple sites. You will manage admission, routing, and autoscaling for high-concurrency workloads, ensuring low latency and zero data retention.

You will collaborate with GPU Performance, Platform, and Backend teams to optimize CUDA workloads, pipelines, and Kubernetes clusters, while delivering robust, scalable serving.

Qualifications

  • 5+ years in distributed systems, ML infra, or production serving; GPU workloads a plus.
  • Hands-on experience running LLMs in production with vLLM, SGLang or TensorRT-LLM.
  • Experience with NVIDIA Dynamo or comparable disaggregated serving (llm-d).
  • Kubernetes in production: deployments, rollouts, operators, or GPU scheduling.
  • Strong fundamentals in queues, backpressure, routing, and burst-load retries.
  • Proficiency in Python and reading systems code; comfortable with AI-assisted tools.

Responsibilities

  • Disaggregated serving: extend Dynamo deployment; manage prefill/decode split and cross-node capacity.
  • Admission & routing: own router and KV-aware routing; implement burst queues and cache-tuning collaboration.
  • Autoscaling: scale prefill/decode to meet latency targets; decide GPU allocation per model pool.

Skills

Distributed systems
GPU workloads
Python
Performance tuning
Cloud/On-prem infra

Education

BSc/MSc in Computer Science, Software Engineering, Electrical Engineering or related field

Tools

Kubernetes
Prometheus
Grafana
Loki

Job description

TensorX, a Dublin-based sovereign AI infrastructure platform, seeks a Senior Inference Engineer to own the disaggregated serving layer across multiple sites. You will manage admission, routing, and autoscaling for high-concurrency workloads, ensuring low latency and zero data retention.

You will collaborate with GPU Performance, Platform, and Backend teams to optimize CUDA workloads, pipelines, and Kubernetes clusters, while delivering robust, scalable serving.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Inference Engineer
Senior Inference Engineer

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working
25 days paid annual leave
Free inference tokens!
GPU Inference Performance Engineer — Hybrid Dublin
GPU Inference Performance Engineer — Hybrid Dublin

TensorX • Dublin

Hybrid
EUR 90,000 - 130,000
25 days paid annual leave
Hybrid working in Dublin
Free inference tokens!
GPU Performance Engineer — Open-Weight AI, Hybrid Dublin
GPU Performance Engineer — Open-Weight AI, Hybrid Dublin

HireHive • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working
25 days paid annual leave
Free inference tokens
Senior Machine Learning Engineer (Inference)
Senior Machine Learning Engineer (Inference)

TensorX • Dublin

On-site
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Senior Infrastructure Engineer (GPU Cloud)
Senior Infrastructure Engineer (GPU Cloud)

Uniting Holding • Dublin

On-site
EUR 75,000 - 95,000
25 days paid annual leave
Free inference tokens
Remote flexibility
Senior ML Systems Engineer (Inference) – Hybrid + Tokens
Senior ML Systems Engineer (Inference) – Hybrid + Tokens

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Engineering Manager, AI Platform & Delivery (Hybrid Dublin)
Engineering Manager, AI Platform & Delivery (Hybrid Dublin)

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Senior Full-Stack Engineer - Hybrid Dublin
Senior Full-Stack Engineer - Hybrid Dublin

TensorX • Dublin

Hybrid
EUR 90,000 - 120,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
GPU Performance Engineer (Inference)
GPU Performance Engineer (Inference)

HireHive • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working
25 days paid annual leave
Free inference tokens
Engineering Manager
Engineering Manager

HireHive • Dublin

On-site
EUR 120,000 - 160,000
Hybrid working from Dublin office
25 days annual leave
Free inference tokens