GPU Performance Engineer — Open-Weight AI, Hybrid Dublin

HireHive

Dublin

Hybrid

EUR 120,000 - 180,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Hybrid working
25 days paid annual leave
Free inference tokens

Job summary

TensorX in Dublin is hiring multiple GPU Performance Engineers to optimize inference engines for latency targets. You will patch open-source engines, tune kernels, and bring up new models on day zero within a high-concurrency production environment.

You will work with the Inference Team on CUDA kernels, memory management and KV cache behavior, striving to maximize per-GPU throughput while maintaining correctness and observability.

Qualifications

  • Ability to read inference engine source to identify root causes.
  • Strong GPU architecture understanding and CUDA fundamentals.
  • Working knowledge of transformer inference and KV cache behavior.
  • Proficient in Python; comfortable reading C++ and CUDA.
  • Measure-first mindset with controlled experiments and timing checks.
  • Honest about results, including failed experiments.
  • Eager to learn; stack evolves weekly with new tech.
  • Familiar with AI-assisted development tools such as Claude Code/Codex.
  • Clear communicator able to explain decisions to technical and non-technical audiences.

Responsibilities

  • Engine tickets end to end: reproduce, patch, test and roll out fixes.
  • Diagnosis across GPU memory, scheduler, container, router and traffic shape.
  • Improve per-GPU goodput by locating bottlenecks and validating gains on production traffic.
  • Patch SGLang and vLLM; upstream fixes when appropriate.
  • Modify and optimize GPU kernels (attention/indexer, FP8/FP4 paths).
  • Model bring‑up on day zero with parallelism layouts and cache config testing.
  • Tune cache behavior and router coordination; document changes.
  • Maintain pre-production gate; provide benchmarks for repeatability.
  • Communicate findings clearly for both engineers and customers.

Skills

GPU architecture
CUDA fundamentals
Transformer inference
Python
C++
Kernel development
KV cache
Benchmarking
AI tooling

Education

BSc/MSc/PhD in CS/Engineering or ML

Tools

SGLang
vLLM
TensorRT-LLM
CUDA

Job description

TensorX in Dublin is hiring multiple GPU Performance Engineers to optimize inference engines for latency targets. You will patch open-source engines, tune kernels, and bring up new models on day zero within a high-concurrency production environment.

You will work with the Inference Team on CUDA kernels, memory management and KV cache behavior, striving to maximize per-GPU throughput while maintaining correctness and observability.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Inference Performance Engineer — Hybrid Dublin
GPU Inference Performance Engineer — Hybrid Dublin

TensorX • Dublin

Hybrid
EUR 90,000 - 130,000
25 days paid annual leave
Hybrid working in Dublin
Free inference tokens!
Senior Inference Engineer — GPU-Disaggregated Serving
Senior Inference Engineer — GPU-Disaggregated Serving

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working
25 days paid annual leave
Free inference tokens!
Cloud AI Performance Engineer
Cloud AI Performance Engineer

Qualcomm • Ireland

On-site
EUR 90,000 - 150,000
Stock bonus
Employee stock purchase scheme
Pension matching scheme
+4
AI Frameworks Engineer - High-Perf DL & OpenVINO
AI Frameworks Engineer - High-Perf DL & OpenVINO

Intel Corporation • Leixlip

Hybrid
EUR 77,000 - 142,000
Hybrid work model
GPU Performance Engineer (Inference)
GPU Performance Engineer (Inference)

HireHive • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working
25 days paid annual leave
Free inference tokens
Senior ML Systems Engineer (Inference) – Hybrid + Tokens
Senior ML Systems Engineer (Inference) – Hybrid + Tokens

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Engineering Manager, AI Platform & Delivery (Hybrid Dublin)
Engineering Manager, AI Platform & Delivery (Hybrid Dublin)

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
AI Inference Performance Engineer
AI Inference Performance Engineer

QUALCOMM, Inc. • Cork

On-site
EUR 90,000 - 150,000
Stock options
Performance bonus
Relocation support
+1
GPU Performance Engineer (Inference)
GPU Performance Engineer (Inference)

TensorX • Dublin

Hybrid
EUR 90,000 - 130,000
25 days paid annual leave
Hybrid working in Dublin
Free inference tokens!
Senior LLM Inference & GPU Performance Engineer
Senior LLM Inference & GPU Performance Engineer

Confidential • Ireland

On-site
EUR 70,000 - 90,000