Senior Inference Engineer

TensorX

Dublin

Hybrid

EUR 120,000 - 180,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Hybrid working
25 days paid annual leave
Free inference tokens!

Job summary

TensorX, a Dublin-based sovereign AI infrastructure platform, seeks a Senior Inference Engineer to own the disaggregated serving layer across multiple sites. You will manage admission, routing, and autoscaling for high-concurrency workloads, ensuring low latency and zero data retention.

You will collaborate with GPU Performance, Platform, and Backend teams to optimize CUDA workloads, pipelines, and Kubernetes clusters, while delivering robust, scalable serving.

Qualifications

  • 5+ years in distributed systems, ML infra, or production serving; GPU workloads a plus.
  • Hands-on experience running LLMs in production with vLLM, SGLang or TensorRT-LLM.
  • Experience with NVIDIA Dynamo or comparable disaggregated serving (llm-d).
  • Kubernetes in production: deployments, rollouts, operators, or GPU scheduling.
  • Strong fundamentals in queues, backpressure, routing, and burst-load retries.
  • Proficiency in Python and reading systems code; comfortable with AI-assisted tools.

Responsibilities

  • Disaggregated serving: extend Dynamo deployment; manage prefill/decode split and cross-node capacity.
  • Admission & routing: own router and KV-aware routing; implement burst queues and cache-tuning collaboration.
  • Autoscaling: scale prefill/decode to meet latency targets; decide GPU allocation per model pool.

Skills

Distributed systems
GPU workloads
Python
Performance tuning
Cloud/On-prem infra

Education

BSc/MSc in Computer Science, Software Engineering, Electrical Engineering or related field

Tools

Kubernetes
Prometheus
Grafana
Loki

Job description

Overview

TensorX is a sovereign AI infrastructure platform headquartered in Dublin. We run frontier open-weight large language models on our own NVIDIA Blackwell GPUs in European datacentres, under EU jurisdiction. Customers reach them through a drop-in OpenAI-compatible API. Nothing they send is retained after the request completes. We serve regulated enterprises in finance, healthcare and government as well as developers and AI platforms. We help them adopt AI without compromising on data privacy, compliance or performance.

We are looking for a Senior Inference Engineer to join our growing engineering team. Reporting to the CTO, you will own the serving layer between our API gateway and the inference engines across more than one site. This covers admission, routing, the prefill/decode split, KV cache movement across the GPU fabric and autoscaling.

Our customers send very long contexts at high concurrency, often in bursts. The number that decides our margin is not how fast one engine runs. It is how many requests the whole fleet answers inside a latency target when a burst arrives. We run NVIDIA Dynamo in production for disaggregated serving on Kubernetes and are moving the rest of the fleet onto it. We are an AI-native team. Tools such as Claude Code and Codex are part of our daily workflow and materially accelerate how we build and operate systems.

You will work side by side with our GPU Performance Team, who own the inside of the engines: kernels, engine patches and per-model tuning. They make each worker faster. You make the fleet of workers behave. Our platform team builds and maintains the Kubernetes clusters and hosts. Our backend team owns the API gateway. You will work with all three every day.

This is a high-impact senior individual contributor role spanning disaggregated serving, routing, autoscaling and multi-site operations.

Responsibilities
  • Disaggregated serving - Run and extend our NVIDIA Dynamo deployment. Own the split between prefill and decode workers. Move capacity between them as the traffic shape changes.

  • Admission & routing - Own the router and KV-aware routing so requests land where their prefix is already cached. Build admission control so a burst queues instead of crashing a worker. The GPU Performance Team tunes cache behaviour inside the engine and contributes to the router code.

  • Autoscaling & capacity - Scale prefill and decode against latency targets and scale down when traffic drops. Decide how many GPUs each model pool needs and where it runs. Base these decisions on traffic data and the parallelism layouts the GPU Performance Team validates.

  • KV cache transfer - Own cross-node KV cache movement (NIXL, Mooncake or similar). Test it across nodes so the network is part of every measurement. Add custom telemetry where the stock tools do not see.

  • Multi-site serving - Run several sites behind one front door, each with its own local serving plane. Keep a path open to non-NVIDIA pools through llm-d when we need one.

  • Observability - Instrument the serving layer with per-model metrics that mean something: time to first token, inter-token latency, queue depth, cache hit rate and errors. Work with Prometheus, Grafana and Loki alongside GPU telemetry.

  • Rollouts & pre-production - Roll out serving-layer changes with zero downtime: drain, verify outside the pool, return, repeat. Share the pre-production gate with the GPU Performance Team so no change reaches a customer unmeasured.

  • Data retention - Make sure nothing a customer sends is retained after the request completes. Enforce this by design in the serving path, not only by policy.

  • Reliability & incidents - Own the serving-layer runbook. Lead the response when the fleet, rather than one engine, misbehaves. Write up each incident so it does not happen twice.

Skills & Experience
  • 5+ years of professional experience in distributed systems, ML infrastructure or production serving. A meaningful portion of this should be on GPU workloads

  • Hands-on experience running large language models in production with vLLM, SGLang or TensorRT-LLM. You understand how prefill, decode, KV cache and batching behave under load

  • Experience with NVIDIA Dynamo or a comparable disaggregated serving system (e.g. llm-d)

  • Kubernetes in production, including deployments, rollouts, operators or controllers and GPU scheduling

  • Strong distributed systems fundamentals: queueing, backpressure, admission control, routing and retries under burst load

  • Experience instrumenting production systems (e.g. Prometheus, Grafana, Loki) and using that data to guide tuning decisions

  • Familiarity with benchmarking methodology: one variable at a time, a pass or fail bar set before the run and results someone else can reproduce

  • Proficiency in Python and comfort reading systems code in other languages

  • Comfortable using AI-assisted development tools (e.g. Claude Code, Codex) as part of your daily workflow

  • A clear and concise communicator who thrives in ambiguity and can articulate technical decisions to both technical and non-technical audiences

Nice to Have
  • Experience with NIXL, Mooncake, LMCache or other KV cache transfer and offload layers

  • Experience debugging performance across a GPU fabric (RDMA, RoCE or InfiniBand)

  • Gateway API Inference Extension or similar Kubernetes-native inference routing

  • Multi-site or multi-cluster operations

  • Contributions to open‑source inference or serving projects

Why This Role
  • Disaggregated serving is live here, not on a roadmap.You will extend a Dynamo deployment that already carries production traffic.

  • The fleet is ours.Our own NVIDIA B300 GPUs in Dublin and Helsinki. You are not renting time on someone else's cluster.

  • The traffic is real.Long-context, bursty, high-concurrency workloads that break assumptions benchmarks never test.

  • The results are real.Our inference stack answers roughly twice as many requests inside the latency target as a standard configuration. We measured this on the same hardware with the same production traffic patterns.

  • You shape the architecture.Dynamo is live, but the multi-site serving layer around it is still being designed. Your decisions become the architecture.

Education & Qualifications
  • BSc/MSc in Computer Science, Software Engineering, Electrical Engineering or a related technical discipline OR equivalent practical experience

Remuneration
  • Highly competitive package, dependent on experience

  • 25 days paid annual leave

  • Hybrid working from our centrally located Dublin office, with remote flexibility

  • Free inference tokens!

NO AGENCY ASSISTANCE REQUIRED
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Performance Engineer (Inference)
GPU Performance Engineer (Inference)

HireHive • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working
25 days paid annual leave
Free inference tokens
GPU Performance Engineer (Inference)
GPU Performance Engineer (Inference)

TensorX • Dublin

Hybrid
EUR 90,000 - 130,000
25 days paid annual leave
Hybrid working in Dublin
Free inference tokens!
Senior Machine Learning Engineer (Inference)
Senior Machine Learning Engineer (Inference)

TensorX • Dublin

On-site
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Senior Infrastructure Engineer (GPU Cloud)
Senior Infrastructure Engineer (GPU Cloud)

Uniting Holding • Dublin

On-site
EUR 75,000 - 95,000
25 days paid annual leave
Free inference tokens
Remote flexibility
Engineering Manager
Engineering Manager

HireHive • Dublin

On-site
EUR 120,000 - 160,000
Hybrid working from Dublin office
25 days annual leave
Free inference tokens
Engineering Manager
Engineering Manager

Uniting Holding • Dublin

On-site
EUR 120,000 - 180,000
Hybrid Dublin office
Free inference tokens
25 days annual leave
Engineering Manager
Engineering Manager

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Senior Inference Engineer — GPU-Disaggregated Serving
Senior Inference Engineer — GPU-Disaggregated Serving

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working
25 days paid annual leave
Free inference tokens!
Senior Full Stack Engineer
Senior Full Stack Engineer

TensorX • Dublin

Hybrid
EUR 90,000 - 120,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Senior Full Stack Engineer
Senior Full Stack Engineer

Uniting Holding • Dublin

On-site
EUR 90,000 - 150,000
Hybrid work in Dublin office
25 days paid annual leave
Free inference tokens!