Get more replies from employers
Send a job-specific resume in minutes.
Uniting Holding seeks a Senior ML Systems Engineer (Inference) in Dublin to lead model serving and performance optimization. This role involves deploying and tuning large language models while directly impacting performance and reliability.
Ideal candidates will have over five years of experience in ML infrastructure, knowledge of GPU architectures, and proficiency in Python. The position offers a flexible hybrid work environment with competitive remuneration.
Tensorix is a sovereign AI infrastructure platform headquartered in Dublin. We deploy and operate open-source large language models on EU-sovereign infrastructure across Europe, providing private, zero-retention inference for regulated industries including finance, healthcare and government. Our platform offers drop-in OpenAI-compatible APIs, enabling developers and enterprises alike to adopt AI without compromising on data privacy, compliance or performance.
We are looking for a Senior ML Systems Engineer (Inference) to join our growing engineering team. Reporting to the CTO, you will act as the technical owner of the model serving layer at the heart of our platform, from selecting and evaluating new open-source models to deploying, tuning and operating them in production on our on-prem GPU fleet. This is a deeply hands‑on role in a fast‑moving scaleup where your work will directly shape the performance, cost and reliability of every token we serve.
You will work primarily with modern inference frameworks such as vLLM, SGLang and TensorRT-LLM, running on NVIDIA hardware across our on‑prem estate, with supporting workloads on AWS. You will benchmark frontier open‑weight models as they release, quantify performance and cost trade‑offs, and lead the technical side of our GPU procurement and capacity planning. We are an AI‑native team - tools such as Claude Code and Codex are part of our daily workflow and materially accelerate how we build and operate systems. We value engineers who combine deep systems intuition with a pragmatic, research‑aware mindset.
This is a high‑impact senior individual contributor role spanning model serving, performance engineering and GPU infrastructure strategy.
5+ years of professional experience (or equivalent depth of expertise) in ML infrastructure, systems engineering or a closely related discipline, with a meaningful portion focused on production ML workloads
Hands‑on experience deploying and tuning large language models with modern inference frameworks such as vLLM, SGLang, TensorRT-LLM and similar high‑performance inference systems
Strong working knowledge of GPU architecture, CUDA fundamentals and the performance characteristics of modern NVIDIA hardware (H100, H200 and B300‑class hardware)
Practical experience with inference optimisation techniques including quantisation (AWQ, GPTQ, FP8), continuous batching, KV cache strategies and tensor/pipeline parallelism
Proficiency in Python and comfort reading and contributing to systems‑level code in the broader inference ecosystem
Solid experience with Linux, containerisation and orchestration of GPU workloads
Familiarity with benchmarking methodology and the ability to design experiments that produce defensible, reproducible results
Comfortable using AI‑assisted development tools (e.g. Claude Code, Codex) as part of your daily workflow
A clear and concise communicator who thrives in ambiguity and can articulate technical decisions to both technical and non‑technical audiences
Experience with Kubernetes and GPU scheduling in multi‑tenant environments
Exposure to distributed training or fine‑tuning workflows, even if your primary focus is inference
Experience with AWS infrastructure and related services (e.g. EC2, ECS, EKS, S3)
Familiarity with alternative accelerators (AMD Instinct, Intel Gaudi) or emerging inference hardware
Contributions to open‑source inference projects such as vLLM, SGLang or related tooling
Exposure to Golang or Rust for systems‑level work
BSc/MSc in Computer Science, Software Engineering, Electrical Engineering OR a related technical discipline OR equivalent practical experience
Highly competitive package, dependent on experience
25 days paid annual leave
Hybrid working from our centrally located Dublin office, with remote flexibility
Free inference tokens!