Senior ML Systems Engineer (Inference) – Hybrid + Tokens

TensorX

Dublin

Hybrid

EUR 120,000 - 180,000

Full time

12 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens

Job summary

Tensorix in Dublin is seeking a Senior ML Systems Engineer (Inference) to own the model serving layer, deploying and tuning open-source LLMs on on-prem GPU fleets and AWS. This hands-on role shapes performance, cost and reliability of token delivery.

You will work with vLLM, SGLang and TensorRT-LLM on NVIDIA hardware, drive hardware planning, benchmarking, and collaborate with platform and product teams. Hybrid work from Dublin with remote flexibility, 25 days leave, and a competitive package.

Qualifications

  • 5+ years of professional experience in ML infrastructure or production ML workloads.
  • Hands-on experience deploying and tuning large language models with modern inference frameworks.
  • Strong knowledge of GPU architecture and NVIDIA hardware (H100/H200).
  • Experience with inference optimisation techniques and benchmarking methodologies.
  • Proficiency in Python and systems-level code reading/writing.
  • Familiarity with Linux, containers and orchestration of GPU workloads.

Responsibilities

  • Deploy and operate open-source LLMs in production using vLLM, SGLang, TensorRT-LLM and similar frameworks.
  • Profile and tune inference workloads for latency, throughput and GPU utilisation.
  • Benchmark new models and maintain internal tooling for reproducible results.
  • Lead GPU procurement planning and capacity planning.
  • Build and maintain GPU infra spanning on-prem and cloud workloads.
  • Instrument serving stack with metrics and drive reliability improvements.
  • Prototype new optimisation techniques and stay current with research.

Skills

ML infrastructure
System engineering
Python
Linux
GPU architecture
CUDA fundamentals
Benchmarking
Communication
Ambiguity tolerance

Education

BSc/MSc in CS or related

Tools

vLLM
SGLang
TensorRT-LLM
Kubernetes
AWS

Job description

Tensorix in Dublin is seeking a Senior ML Systems Engineer (Inference) to own the model serving layer, deploying and tuning open-source LLMs on on-prem GPU fleets and AWS. This hands-on role shapes performance, cost and reliability of token delivery.

You will work with vLLM, SGLang and TensorRT-LLM on NVIDIA hardware, drive hardware planning, benchmarking, and collaborate with platform and product teams. Hybrid work from Dublin with remote flexibility, 25 days leave, and a competitive package.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer (Inference)
Senior ML Systems Engineer (Inference)

Uniting Holding • Dublin

Hybrid
EUR 80,000 - 120,000
25 days paid annual leave
Free inference tokens
Remote work flexibility
Senior ML Inference Systems Engineer (Hybrid, On-Prem GPU)
Senior ML Inference Systems Engineer (Hybrid, On-Prem GPU)

Uniting Holding • Dublin

Hybrid
EUR 80,000 - 120,000
25 days paid annual leave
Free inference tokens
Remote work flexibility
Senior Machine Learning Engineer (Inference)
Senior Machine Learning Engineer (Inference)

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Hybrid/Remote Senior AI Inference Engineer
Hybrid/Remote Senior AI Inference Engineer

F5 Networks, Inc.  • Dublin

Hybrid
EUR 120,000 - 180,000
Senior ML Engineer - NLP, Production-Ready - Hybrid
Senior ML Engineer - NLP, Production-Ready - Hybrid

Docusign • Dublin

Hybrid
EUR 80,000 - 100,000
Senior AI Engineer, Payments & FinTech (Hybrid, Dublin)
Senior AI Engineer, Payments & FinTech (Hybrid, Dublin)

Intellect • Dublin

Hybrid
EUR 120,000 - 180,000
Senior LLM Inference & GPU Performance Engineer
Senior LLM Inference & GPU Performance Engineer

Confidential • Ireland

On-site
EUR 70,000 - 90,000
AI Inference Engineer — High-Performance, Low-Latency ML
AI Inference Engineer — High-Performance, Low-Latency ML

F5 • Dublin

On-site
EUR 90,000 - 150,000
Senior Full-Stack Engineer — Hybrid Dublin
Senior Full-Stack Engineer — Hybrid Dublin

TensorX • Dublin

Hybrid
EUR 90,000 - 120,000
25 days paid annual leave
Hybrid working from Dublin office
Free inference tokens
Senior ML Software Engineer - Hybrid & Impact
Senior ML Software Engineer - Hybrid & Impact

Arm • Galway

Hybrid
EUR 97,000 - 133,000
Training and professional development
Friendly working environment
Flexibility in hybrid work