GPU Inference Performance Engineer — Hybrid Dublin

TensorX

Dublin

Hybrid

EUR 90,000 - 130,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

25 days paid annual leave
Hybrid working in Dublin
Free inference tokens!

Job summary

TensorX, based in Dublin, is seeking GPU Performance Engineers to optimize inference engines, patch open-source components, and push improvements across a live fleet. You will work on kernel-level performance, model bring-up, and cache behavior to meet tight latency targets in a high-concurrency environment.

Reporting to the CTO, you’ll patch defects in SGLang and related tools, verify fixes on production traffic, and collaborate with the Inference Team to size pools and improve throughput.

Qualifications

  • Read inference engine source to identify mechanism behind problems.
  • Solid understanding of GPU architecture and CUDA fundamentals.
  • Working knowledge of transformer inference, attention variants, KV cache and quantisation.
  • Proficiency in Python and comfortable reading C++/CUDA.
  • Measure and report results with clear trade-offs; test timing and correctness.

Responsibilities

  • Engine tickets end to end: reproduce, diagnose, patch, verify, and roll out fixes.
  • Investigation & diagnosis across GPU memory, engine scheduler, container, router, and traffic shape.
  • Improve Goodput per GPU by locating bottlenecks in kernels, scheduler, or configuration.
  • Patch SGLang and upstream defects; keep image patch sets consistent across fleet.
  • Bring up new models on day zero; test layouts and caching strategies.
  • Document findings with repeatable benchmarks for the team and customers.

Skills

GPU architecture
CUDA fundamentals
Transformer inference
Python
C++ & CUDA
Problem solving
Communication
Open-source patching

Education

BSc/MSc/PhD in CS/Engineering/ML

Tools

SGLang
vLLM
TensorRT-LLM

Job description

TensorX, based in Dublin, is seeking GPU Performance Engineers to optimize inference engines, patch open-source components, and push improvements across a live fleet. You will work on kernel-level performance, model bring-up, and cache behavior to meet tight latency targets in a high-concurrency environment.

Reporting to the CTO, you’ll patch defects in SGLang and related tools, verify fixes on production traffic, and collaborate with the Inference Team to size pools and improve throughput.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Performance Engineer — Open-Weight AI, Hybrid Dublin
GPU Performance Engineer — Open-Weight AI, Hybrid Dublin

HireHive • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working
25 days paid annual leave
Free inference tokens
Senior Inference Engineer — GPU-Disaggregated Serving
Senior Inference Engineer — GPU-Disaggregated Serving

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working
25 days paid annual leave
Free inference tokens!
GPU Performance Engineer (Inference)
GPU Performance Engineer (Inference)

HireHive • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working
25 days paid annual leave
Free inference tokens
GPU Performance Engineer (Inference)
GPU Performance Engineer (Inference)

TensorX • Dublin

Hybrid
EUR 90,000 - 130,000
25 days paid annual leave
Hybrid working in Dublin
Free inference tokens!
Senior LLM Inference & GPU Performance Engineer
Senior LLM Inference & GPU Performance Engineer

Confidential • Ireland

On-site
EUR 70,000 - 90,000
Senior ML Systems Engineer (Inference) – Hybrid + Tokens
Senior ML Systems Engineer (Inference) – Hybrid + Tokens

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Hybrid Senior GPU Cloud Infra Engineer (Dublin)
Hybrid Senior GPU Cloud Infra Engineer (Dublin)

Uniting Holding • Dublin

Hybrid
EUR 75,000 - 95,000
25 days paid annual leave
Free inference tokens
Remote flexibility
Engineering Manager, AI Platform & Delivery (Hybrid Dublin)
Engineering Manager, AI Platform & Delivery (Hybrid Dublin)

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Senior Machine Learning Engineer (Inference)
Senior Machine Learning Engineer (Inference)

TensorX • Dublin

On-site
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Senior GPU Architect & Performance Engineer
Senior GPU Architect & Performance Engineer

Qualcomm • Cork

On-site
EUR 90,000 - 140,000
Maternity/Paternity Leave
Employee stock purchase scheme
Matching pension scheme
+5