GPU Performance Engineer (Inference)

HireHive

Dublin

Hybrid

EUR 120,000 - 180,000

Full time

9 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Hybrid working
25 days paid annual leave
Free inference tokens

Job summary

TensorX in Dublin is hiring multiple GPU Performance Engineers to optimize inference engines for latency targets. You will patch open-source engines, tune kernels, and bring up new models on day zero within a high-concurrency production environment.

You will work with the Inference Team on CUDA kernels, memory management and KV cache behavior, striving to maximize per-GPU throughput while maintaining correctness and observability.

Qualifications

  • Ability to read inference engine source to identify root causes.
  • Strong GPU architecture understanding and CUDA fundamentals.
  • Working knowledge of transformer inference and KV cache behavior.
  • Proficient in Python; comfortable reading C++ and CUDA.
  • Measure-first mindset with controlled experiments and timing checks.
  • Honest about results, including failed experiments.
  • Eager to learn; stack evolves weekly with new tech.
  • Familiar with AI-assisted development tools such as Claude Code/Codex.
  • Clear communicator able to explain decisions to technical and non-technical audiences.

Responsibilities

  • Engine tickets end to end: reproduce, patch, test and roll out fixes.
  • Diagnosis across GPU memory, scheduler, container, router and traffic shape.
  • Improve per-GPU goodput by locating bottlenecks and validating gains on production traffic.
  • Patch SGLang and vLLM; upstream fixes when appropriate.
  • Modify and optimize GPU kernels (attention/indexer, FP8/FP4 paths).
  • Model bring‑up on day zero with parallelism layouts and cache config testing.
  • Tune cache behavior and router coordination; document changes.
  • Maintain pre-production gate; provide benchmarks for repeatability.
  • Communicate findings clearly for both engineers and customers.

Skills

GPU architecture
CUDA fundamentals
Transformer inference
Python
C++
Kernel development
KV cache
Benchmarking
AI tooling

Education

BSc/MSc/PhD in CS/Engineering or ML

Tools

SGLang
vLLM
TensorRT-LLM
CUDA

Job description

Overview

TensorX is a sovereign AI infrastructure platform headquartered in Dublin. We run frontier open-weight large language models on our own NVIDIA Blackwell GPUs in European datacentres, under EU jurisdiction. Customers reach them through a drop-in OpenAI-compatible API. Nothing they send is retained after the request completes. We serve regulated enterprises in finance, healthcare and government as well as developers and AI platforms. We help them adopt AI without compromising on data privacy, compliance or performance.

We are looking for multiple GPU Performance Engineers to join our growing engineering team. Reporting to the CTO, you will make each inference engine in our fleet do more work inside its latency target. You will take engine problems from report to fix, patch the open-source engines we depend on and go beneath them to the kernel when the engine is the limit.

Our margin depends on one number: how many requests each GPU answers inside a latency target. That is not the same as peak tokens per second on a benchmark. Passengers carried, not miles per hour. Our customers send very long contexts at high concurrency, so the constraint that matters is rarely the one the headline number measures.

The serving engines we depend on, SGLang first, are open source and still maturing. We hit their defects early because we run new models on day zero. We patch them, tune them and change the kernel where the engine itself is the limit. A new open-weight model worth serving arrives most weeks and more GPU capacity is going live now. We are an AI-native team. Tools such as Claude Code and Codex are part of our daily workflow and materially accelerate how we build and operate systems.

You will work side by side with our Inference Team, who run NVIDIA Dynamo and the serving fleet. You make each worker faster. They make the fleet of workers behave.

We hire on evidence of skill, not years of experience. We will hire across multiple levels. One of the seats may suit someone earlier in their career.

This is a hands‑on individual contributor role spanning engine performance, GPU kernels, model bring-up and cache behaviour.

Responsibilities
  • Engine tickets end to end - Take engine problems from report to fix. A model runs out of memory at long context, a patch cuts throughput or a new release breaks tool calling. Reproduce it, find the mechanism, test a fix, patch it and roll it out. Then write down what you found.

  • Investigation & diagnosis - Debug across GPU memory, the engine scheduler, the container, the router and the traffic shape. Many of our hardest problems sit where two of these meet. When you raise a problem with the team, bring a read, not a question: what you think it is and why.

  • Goodput per GPU - Goodput is the number of requests each GPU answers inside the latency target. Find where it is being lost, whether in a kernel, the scheduler, the router or a configuration flag. Win it back. Measure every gain on production traffic patterns rather than a synthetic load.

  • Engine patches & upstream - Patch SGLang and vLLM defects on new models. Bake each fix into the image and keep the patch set consistent across the fleet. Where a fix is not specific to us, open the pull request upstream, not just the issue.

  • GPU kernels - When profiling shows an engine kernel is the bottleneck, write or modify it. Examples include Blackwell attention/indexer paths, FP8/FP4 paths and memory-bound decode. Profile before you form an opinion.

  • Model bring‑up - Bring up new models on day zero. Run A/B bake‑offs across parallelism layouts, KV cache configurations and speculative decoding settings. Test tensor, data and expert parallelism for each model. Give the Inference Team the best layout for each model so they can size the pools.

  • Cache behaviour - Tune prefix caching, KV cache behaviour and KV replication under tensor parallelism inside the engine. Work with the Inference Team on cache‑aware routing and contribute to the router code.

  • Pre‑production gate - Share the pre‑production gate with the Inference Team so no change reaches a customer unmeasured. Run it on your own changes and on every model you bring up.

  • Documentation - Write up every finding as a report someone else can rerun. Explain trade‑offs plainly enough for the Inference Team and for customers. Keep a benchmark method that works without you in the room.

Skills & Experience
  • Evidence that you can read an inference engine's source to find the mechanism behind a problem rather than only the symptom. A pull request, a write-up or a post‑mortem we can check

  • Solid understanding of GPU architecture and CUDA fundamentals: memory hierarchy, occupancy and what makes a kernel compute‑bound or memory‑bound

  • Working knowledge of transformer inference: attention variants (MLA, DSA, GQA), KV cache behaviour, continuous batching and the accuracy impact of quantisation

  • Proficiency in Python and comfort reading C++ and CUDA

  • A measure‑first habit: one variable per test arm, a pass or fail bar set before you run and a correctness check before you trust a timing

  • Honest about results, including the ones that did not work out. Losing to a baseline and saying so counts in your favour

  • A drive to learn. The stack changes every week. We value someone who reads the source over someone who already knows last year's answer

  • Comfortable using AI‑assisted development tools (e.g. Claude Code, Codex) as part of your daily workflow

  • A clear and concise communicator who thrives in ambiguity and can articulate technical decisions to both technical and non-technical audiences

Nice to Have
  • A track record of writing and benchmarking CUDA kernels on Hopper or Blackwell

  • Upstream contributions to SGLang, vLLM or TensorRT‑LLM

  • Experience with Triton, CUTLASS, CuTe, ThunderKittens or similar kernel tools

  • Kernel entries in GPU Mode leaderboards, MLSys contests or FlashInfer challenges

  • Experience profiling memory behaviour and out‑of‑memory errors on large models

  • Hands‑on time with Blackwell features (tcgen05, TMEM, TMA)

  • Technical writing in public

Why This Role
  • You work inside the engine. Many GPU clouds hire performance engineers for cold starts, storage and containers. We hire them to change the engine and the kernel. That is where our margin is.

  • The fleet is ours. Our own NVIDIA B300 GPUs in Dublin and Helsinki. You are not renting time on someone else's cluster.

  • The traffic is real. Long-context, high-concurrency production workloads that break assumptions benchmarks never test.

  • The results are real. Our inference stack answers roughly twice as many requests inside the latency target as a standard configuration. We measured this on the same hardware with the same production traffic patterns.

  • The engine is open. We patch SGLang in production and send the fixes upstream. If there is a paper in the work, we would rather it went out with your name on it.

  • The research list is long. The better the day‑to‑day is covered, the more time goes on KV cache beyond GPU memory, context parallelism, wide expert parallelism and attention/FFN disaggregation.

Education & Qualifications
  • BSc/MSc/PhD in Computer Science, Engineering, Machine Learning or a related technical discipline OR equivalent demonstrable ability

Remuneration
  • Highly competitive package, dependent on experience

  • 25 days paid annual leave

  • Hybrid working from our centrally located Dublin office, with remote flexibility

  • Free inference tokens!

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Performance Engineer (Inference)
GPU Performance Engineer (Inference)

TensorX • Dublin

Hybrid
EUR 90,000 - 130,000
25 days paid annual leave
Hybrid working in Dublin
Free inference tokens!
Senior Inference Engineer
Senior Inference Engineer

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working
25 days paid annual leave
Free inference tokens!
Senior Machine Learning Engineer (Inference)
Senior Machine Learning Engineer (Inference)

TensorX • Dublin

On-site
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Senior Infrastructure Engineer (GPU Cloud)
Senior Infrastructure Engineer (GPU Cloud)

Uniting Holding • Dublin

On-site
EUR 75,000 - 95,000
25 days paid annual leave
Free inference tokens
Remote flexibility
Engineering Manager
Engineering Manager

HireHive • Dublin

On-site
EUR 120,000 - 160,000
Hybrid working from Dublin office
25 days annual leave
Free inference tokens
Engineering Manager
Engineering Manager

Uniting Holding • Dublin

On-site
EUR 120,000 - 180,000
Hybrid Dublin office
Free inference tokens
25 days annual leave
Engineering Manager
Engineering Manager

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
GPU Inference Performance Engineer — Hybrid Dublin
GPU Inference Performance Engineer — Hybrid Dublin

TensorX • Dublin

Hybrid
EUR 90,000 - 130,000
25 days paid annual leave
Hybrid working in Dublin
Free inference tokens!
Senior Full Stack Engineer
Senior Full Stack Engineer

TensorX • Dublin

Hybrid
EUR 90,000 - 120,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Senior Full Stack Engineer
Senior Full Stack Engineer

Uniting Holding • Dublin

On-site
EUR 90,000 - 150,000
Hybrid work in Dublin office
25 days paid annual leave
Free inference tokens!