GPU Performance Engineer (Inference)

TensorX

Dublin

Hybrid

EUR 90,000 - 130,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

25 days paid annual leave
Hybrid working in Dublin
Free inference tokens!

Job summary

TensorX, based in Dublin, is seeking GPU Performance Engineers to optimize inference engines, patch open-source components, and push improvements across a live fleet. You will work on kernel-level performance, model bring-up, and cache behavior to meet tight latency targets in a high-concurrency environment.

Reporting to the CTO, you’ll patch defects in SGLang and related tools, verify fixes on production traffic, and collaborate with the Inference Team to size pools and improve throughput.

Qualifications

  • Read inference engine source to identify mechanism behind problems.
  • Solid understanding of GPU architecture and CUDA fundamentals.
  • Working knowledge of transformer inference, attention variants, KV cache and quantisation.
  • Proficiency in Python and comfortable reading C++/CUDA.
  • Measure and report results with clear trade-offs; test timing and correctness.

Responsibilities

  • Engine tickets end to end: reproduce, diagnose, patch, verify, and roll out fixes.
  • Investigation & diagnosis across GPU memory, engine scheduler, container, router, and traffic shape.
  • Improve Goodput per GPU by locating bottlenecks in kernels, scheduler, or configuration.
  • Patch SGLang and upstream defects; keep image patch sets consistent across fleet.
  • Bring up new models on day zero; test layouts and caching strategies.
  • Document findings with repeatable benchmarks for the team and customers.

Skills

GPU architecture
CUDA fundamentals
Transformer inference
Python
C++ & CUDA
Problem solving
Communication
Open-source patching

Education

BSc/MSc/PhD in CS/Engineering/ML

Tools

SGLang
vLLM
TensorRT-LLM

Job description

Overview

TensorX is a sovereign AI infrastructure platform headquartered in Dublin. We run frontier open-weight large language models on our own NVIDIA Blackwell GPUs in European datacentres, under EU jurisdiction. Customers reach them through a drop-in OpenAI-compatible API. Nothing they send is retained after the request completes. We serve regulated enterprises in finance, healthcare and government as well as developers and AI platforms. We help them adopt AI without compromising on data privacy, compliance or performance.


We are looking for multiple GPU Performance Engineers to join our growing engineering team. Reporting to the CTO, you will make each inference engine in our fleet do more work inside its latency target. You will take engine problems from report to fix, patch the open-source engines we depend on and go beneath them to the kernel when the engine is the limit.


Our margin depends on one number: how many requests each GPU answers inside a latency target. That is not the same as peak tokens per second on a benchmark. Passengers carried, not miles per hour. Our customers send very long contexts at high concurrency, so the constraint that matters is rarely the one the headline number measures.


The serving engines we depend on, SGLang first, are open source and still maturing. We hit their defects early because we run new models on day zero. We patch them, tune them and change the kernel where the engine itself is the limit. A new open-weight model worth serving arrives most weeks and more GPU capacity is going live now. We are an AI-native team. Tools such as Claude Code and Codex are part of our daily workflow and materially accelerate how we build and operate systems.


You will work side by side with our Inference Team, who run NVIDIA Dynamo and the serving fleet. You make each worker faster. They make the fleet of workers behave.


We hire on evidence of skill, not years of experience. We will hire across multiple levels. One of the seats may suit someone earlier in their career.


This is a hands-on individual contributor role spanning engine performance, GPU kernels, model bring-up and cache behaviour.


Responsibilities


  • Engine tickets end to end - Take engine problems from report to fix. A model runs out of memory at long context, a patch cuts throughput or a new release breaks tool calling. Reproduce it, find the mechanism, test a fix, patch it and roll it out. Then write down what you found.


  • Investigation & diagnosis - Debug across GPU memory, the engine scheduler, the container, the router and the traffic shape. Many of our hardest problems sit where two of these meet. When you raise a problem with the team, bring a read, not a question: what you think it is and why.


  • Goodput per GPU - Goodput is the number of requests each GPU answers inside the latency target. Find where it is being lost, whether in a kernel, the scheduler, the router or a configuration flag. Win it back. Measure every gain on production traffic patterns rather than a synthetic load.


  • Engine patches & upstream - Patch SGLang and vLLM defects on new models. Bake each fix into the image and keep the patch set consistent across the fleet. Where a fix is not specific to us, open the pull request upstream, not just the issue.


  • GPU kernels - When profiling shows an engine kernel is the bottleneck, write or modify it. Examples include Blackwell attention/indexer paths, FP8/FP4 paths and memory-bound decode. Profile before you form an opinion.


  • Model bring-up - Bring up new models on day zero. Run A/B bake-offs across parallelism layouts, KV cache configurations and speculative decoding settings. Test tensor, data and expert parallelism for each model. Give the Inference Team the best layout for each model so they can size the pools.


  • Cache behaviour - Tune prefix caching, KV cache behaviour and KV replication under tensor parallelism inside the engine. Work with the Inference Team on cache-aware routing and contribute to the router code.


  • Pre-production gate - Share the pre-production gate with the Inference Team so no change reaches a customer unmeasured. Run it on your own changes and on every model you bring up.


  • Documentation - Write up every finding as a report someone else can rerun. Explain trade-offs plainly enough for the Inference Team and for customers. Keep a benchmark method that works without you in the room.



Skills & Experience


  • Evidence that you can read an inference engine's source to find the mechanism behind a problem rather than only the symptom. A pull request, a write-up or a post-mortem we can check


  • Solid understanding of GPU architecture and CUDA fundamentals: memory hierarchy, occupancy and what makes a kernel compute-bound or memory-bound


  • Working knowledge of transformer inference: attention variants (MLA, DSA, GQA), KV cache behaviour, continuous batching and the accuracy impact of quantisation


  • Proficiency in Python and comfort reading C++ and CUDA


  • A measure-first habit: one variable per test arm, a pass or fail bar set before you run and a correctness check before you trust a timing


  • Honest about results, including the ones that did not work out. Losing to a baseline and saying so counts in your favour


  • A drive to learn. The stack changes every week. We value someone who reads the source over someone who already knows last year's answer


  • Comfortable using AI-assisted development tools (e.g. Claude Code, Codex) as part of your daily workflow


  • A clear and concise communicator who thrives in ambiguity and can articulate technical decisions to both technical and non-technical audiences



Nice to Have


  • A track record of writing and benchmarking CUDA kernels on Hopper or Blackwell


  • Upstream contributions to SGLang, vLLM or TensorRT-LLM


  • Experience with Triton, CUTLASS, CuTe, ThunderKittens or similar kernel tools


  • Kernel entries in GPU Mode leaderboards, MLSys contests or FlashInfer challenges


  • Experience profiling memory behaviour and out-of-memory errors on large models


  • Hands-on time with Blackwell features (tcgen05, TMEM, TMA)


  • Technical writing in public



Why This Role


  • You work inside the engine. Many GPU clouds hire performance engineers for cold starts, storage and containers. We hire them to change the engine and the kernel. That is where our margin is.


  • The fleet is ours. Our own NVIDIA B300 GPUs in Dublin and Helsinki. You are not renting time on someone else’s cluster.


  • The traffic is real. Long-context, high-concurrency production workloads that break assumptions benchmarks never test.


  • The results are real. Our inference stack answers roughly twice as many requests inside the latency target as a standard configuration. We measured this on the same hardware with the same production traffic patterns.


  • The engine is open. We patch SGLang in production and send the fixes upstream. If there is a paper in the work, we would rather it went out with your name on it.


  • The research list is long. The better the day-to-day is covered, the more time goes on KV cache beyond GPU memory, context parallelism, wide expert parallelism and attention/FFN disaggregation.



Education & Qualifications


  • BSc/MSc/PhD in Computer Science, Engineering, Machine Learning or a related technical discipline OR equivalent demonstrable ability



Remuneration


  • Highly competitive package, dependent on experience


  • 25 days paid annual leave


  • Hybrid working from our centrally located Dublin office, with remote flexibility


  • Free inference tokens!


Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Performance Engineer (Inference)
GPU Performance Engineer (Inference)

HireHive • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working
25 days paid annual leave
Free inference tokens
Senior Inference Engineer
Senior Inference Engineer

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working
25 days paid annual leave
Free inference tokens!
Senior Machine Learning Engineer (Inference)
Senior Machine Learning Engineer (Inference)

TensorX • Dublin

On-site
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Senior Infrastructure Engineer (GPU Cloud)
Senior Infrastructure Engineer (GPU Cloud)

Uniting Holding • Dublin

On-site
EUR 75,000 - 95,000
25 days paid annual leave
Free inference tokens
Remote flexibility
Engineering Manager
Engineering Manager

HireHive • Dublin

On-site
EUR 120,000 - 160,000
Hybrid working from Dublin office
25 days annual leave
Free inference tokens
Senior Machine Learning Engineer, LLM Inference Optimization
Senior Machine Learning Engineer, LLM Inference Optimization

Lever, Inc. • Ireland

On-site
EUR 120,000 - 180,000
Competitive pay
Career growth
Ownership over technical work
+2
Engineering Manager
Engineering Manager

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Engineering Manager
Engineering Manager

Uniting Holding • Dublin

On-site
EUR 120,000 - 180,000
Hybrid Dublin office
Free inference tokens
25 days annual leave
GPU Inference Performance Engineer — Hybrid Dublin
GPU Inference Performance Engineer — Hybrid Dublin

TensorX • Dublin

Hybrid
EUR 90,000 - 130,000
25 days paid annual leave
Hybrid working in Dublin
Free inference tokens!
Technical Support Manager
Technical Support Manager

TensorX • Dublin

Hybrid
EUR 85,000 - 120,000
25 days paid annual leave
Hybrid working from Dublin office
Free inference tokens