Senior LLM Inference & GPU Performance Engineer

Confidential

Ireland

On-site

EUR 70,000 - 90,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Confidential company in Ireland is seeking a specialist to optimize the performance of large language models in production, focusing on latency, throughput, and cost. The role involves working deeply at the GPU level and applying various optimization strategies to improve efficiency.

Candidates must have extensive experience in optimizing deep-learning models and hands-on GPU programming skills. If you have a knack for measurable performance improvements and want to work in a high-impact area, this could be for you!

Qualifications

  • Deep experience optimizing deep-learning inference in production.
  • Hands-on GPU programming and performance engineering (CUDA or equivalent).
  • Fluency with modern LLM serving stacks.
  • Track record of measurable performance wins.

Responsibilities

  • Optimize LLM inference for latency, throughput, and cost.
  • Profile and tune GPU performance.
  • Get the most out of serving frameworks.
  • Partner with model and platform teams.

Skills

Optimizing deep-learning inference
GPU programming (CUDA)
Performance engineering
LLM serving stacks (vLLM, TensorRT-LLM)

Job description

Confidential company in Ireland is seeking a specialist to optimize the performance of large language models in production, focusing on latency, throughput, and cost. The role involves working deeply at the GPU level and applying various optimization strategies to improve efficiency.

Candidates must have extensive experience in optimizing deep-learning models and hands-on GPU programming skills. If you have a knack for measurable performance improvements and want to work in a high-impact area, this could be for you!

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Engineer - High-Throughput LLM Serving
Senior AI Inference Engineer - High-Throughput LLM Serving

Confidential • Ireland

On-site
EUR 120,000 - 180,000
Senior ML Inference Systems Engineer (Hybrid, On-Prem GPU)
Senior ML Inference Systems Engineer (Hybrid, On-Prem GPU)

Uniting Holding • Dublin

Hybrid
EUR 80,000 - 120,000
25 days paid annual leave
Free inference tokens
Remote work flexibility
LLM Inference Engineer — Low-Latency, High-Throughput
LLM Inference Engineer — Low-Latency, High-Throughput

F5 Networks, Inc.  • Dublin

On-site
EUR 70,000 - 90,000
Flexible work conditions
Equal employment opportunities
GenAI/LLM Engineer – On-Prem GPU, Remote Ireland
GenAI/LLM Engineer – On-Prem GPU, Remote Ireland

Ergo IT Recruitment Services • Dublin

On-site
EUR 90,000 - 120,000
Senior AI SRE: LLM Infra & GPU-Powered Ops
Senior AI SRE: LLM Infra & GPU-Powered Ops

Confidential • Dublin

On-site
EUR 120,000 - 180,000
Senior AI Model Optimization Architect for Inference
Senior AI Model Optimization Architect for Inference

Qualcomm • Ireland

On-site
EUR 120,000 - 180,000
Salary, stock and performance related—
Relocation and immigration support
Education Assistance
+1
AI Model Optimization Architect for LLMs & Multimodal
AI Model Optimization Architect for LLMs & Multimodal

Qualcomm • Cork

Hybrid
EUR 120,000 - 180,000
Salary and equity package
Relocation support
Education Assistance
+3
Senior ML Systems Engineer (Inference)
Senior ML Systems Engineer (Inference)

Uniting Holding • Dublin

Hybrid
EUR 80,000 - 120,000
25 days paid annual leave
Free inference tokens
Remote work flexibility
LLM Serving Engineer, Cloud AI Platform
LLM Serving Engineer, Cloud AI Platform

Qualcomm • Cork

Hybrid
EUR 110,000 - 160,000
Salary and stock options
Performance bonus
Maternity/Paternity Leave
+8
LLM Inference Engineer — Cloud AI Platform
LLM Inference Engineer — Cloud AI Platform

Qualcomm • Ireland

On-site
EUR 110,000 - 170,000
Salary and bonus
Parental leave
Employee stock purchase
+4