Senior LLM Inference Optimization Engineer

Lever, Inc.

Ireland

On-site

EUR 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Competitive pay
Career growth
Ownership over technical work
Collaborative environment
International teams

Job summary

Lever, Inc. is seeking a Senior Machine Learning Engineer specialized in LLM and Vision-Language Model inference optimization, based in Ireland.

The role focuses on optimizing latency, throughput, memory usage, and cost per token across model artifacts to production deployments. You will work across model internals, serving architectures, and benchmarking, collaborating with kernel, platform, and research teams to deliver measurable production improvements.

Qualifications

  • Strong software engineering skills in Python and PyTorch.
  • Hands-on experience deploying, operating, or optimizing LLM/VLM inference systems.
  • Experience with modern inference stacks such as vLLM, Triton, or TensorRT-LLM.

Responsibilities

  • Own optimization initiatives for specific model families and endpoints.
  • Evaluate inference engines and recommend serving configurations.
  • Diagnose and resolve production performance regressions.
  • Deploy, benchmark, and extend inference engines; optimize latency and throughput.

Skills

Python
PyTorch
LLM Inference
Performance optimization
Benchmarking
CUDA knowledge
Systems deployment
Cross-team collaboration

Tools

vLLM
SGLang
TensorRT-LLM
Triton Inference Server
Ray Serve
KServe
NVIDIA Dynamo

Job description

Lever, Inc. is seeking a Senior Machine Learning Engineer specialized in LLM and Vision-Language Model inference optimization, based in Ireland.

The role focuses on optimizing latency, throughput, memory usage, and cost per token across model artifacts to production deployments. You will work across model internals, serving architectures, and benchmarking, collaborating with kernel, platform, and research teams to deliver measurable production improvements.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior LLM Inference & GPU Performance Engineer
Senior LLM Inference & GPU Performance Engineer

Confidential • Ireland

On-site
EUR 70,000 - 90,000
LLM Inference Engineer — Low-Latency, High-Throughput
LLM Inference Engineer — Low-Latency, High-Throughput

F5 Networks, Inc.  • Dublin

On-site
EUR 70,000 - 90,000
Flexible work conditions
Equal employment opportunities
Senior Machine Learning Engineer, LLM Inference Optimization
Senior Machine Learning Engineer, LLM Inference Optimization

Lever, Inc. • Ireland

On-site
EUR 120,000 - 180,000
Competitive pay
Career growth
Ownership over technical work
+2
AI Model Optimization Architect for LLMs & Multimodal
AI Model Optimization Architect for LLMs & Multimodal

Qualcomm • Cork

Hybrid
EUR 120,000 - 180,000
Salary and equity package
Relocation support
Education Assistance
+3
LLM Inference Engineer — Cloud AI Platform
LLM Inference Engineer — Cloud AI Platform

Qualcomm • Ireland

On-site
EUR 110,000 - 170,000
Salary and bonus
Parental leave
Employee stock purchase
+4
Senior AI Model Optimization Architect for Inference
Senior AI Model Optimization Architect for Inference

Qualcomm • Ireland

On-site
EUR 120,000 - 180,000
Salary, stock and performance related—
Relocation and immigration support
Education Assistance
+1
Senior AI Infra Engineer – LLM Training & Scaling
Senior AI Infra Engineer – LLM Training & Scaling

Jobgether • Ireland

On-site
EUR 90,000 - 120,000
Competitive compensation
Career growth and learning
Ownership over your work
+2
Senior ML Systems Engineer: Arm Acceleration & LLM
Senior ML Systems Engineer: Arm Acceleration & LLM

NLP PEOPLE • Galway

On-site
EUR 120,000 - 160,000
Senior ML Systems Engineer (Inference) – Hybrid + Tokens
Senior ML Systems Engineer (Inference) – Hybrid + Tokens

TensorX • Dublin

Hybrid
EUR 120,000 - 180,000
Hybrid working from Dublin office
25 days paid annual leave
Free inference tokens
Staff AI Model Optimization Architect — Scalable Inference
Staff AI Model Optimization Architect — Scalable Inference

Qualcomm • Cork

On-site
EUR 150,000 - 190,000
Salary and stock bonus
Relocation assistance
Education assistance
+4