Senior AI Inference Performance Engineer - vLLM on GPUs

Advanced Micro Devices

Helsinki

On-site

EUR 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

AMD benefits

Job summary

Advanced Micro Devices in Helsinki, Finland, is seeking a senior AI performance engineer to push inference efficiency on GPUs using vLLM. You will own end-to-end optimization, profiling, and kernel-level improvements across multi-model deployments.

The role requires deep knowledge of vLLM internals, ROCm integration, and end-to-end workload bottleneck diagnosis, with strong Python and C++ skills. Helsinki or Stockholm options are available.

Qualifications

  • 5+ years of software development experience in GPU computing, AI systems, or HPC.
  • Experience with vLLM internals and ROCm integration.
  • Ability to profile end-to-end workloads from user request to GPU kernel.

Responsibilities

  • Drive performance optimization end-to-end on vLLM across leading models and configurations.
  • Profile, diagnose, and resolve cross-stack performance bottlenecks in vLLM deployments.
  • Diagnose kernel-level performance issues and translate findings into actionable optimizations.
  • Contribute to customer-facing engagements with measurable uplifts.
  • Integrate and optimize custom kernels within vLLM and related frameworks.
  • Optimize multi-node distributed inference with scale-out strategies.
  • Contribute to shared performance optimization methodology.
  • Leverage AI agents to accelerate daily work and define best practices.
  • Upstream optimizations into vLLM and adjacent open-source projects.

Skills

vLLM internals
GPU profiling
Python
C++
ROCm
CUDA
SGLang/TensorRT-LLM familiarity
Multi-node distributed inference

Education

Master's or PhD in Computer Science/Computer Engineering/Electrical Engineering

Tools

HIP
CUDA
Triton
CK

Job description

Advanced Micro Devices in Helsinki, Finland, is seeking a senior AI performance engineer to push inference efficiency on GPUs using vLLM. You will own end-to-end optimization, profiling, and kernel-level improvements across multi-model deployments.

The role requires deep knowledge of vLLM internals, ROCm integration, and end-to-end workload bottleneck diagnosis, with strong Python and C++ skills. Helsinki or Stockholm options are available.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Performance Engineer - LLM Inference (vLLM)
Senior AI Performance Engineer - LLM Inference (vLLM)

AMD • Helsinki

Hybrid
EUR 140,000 - 200,000
Senior AI Inference Performance Engineer (LLM)
Senior AI Inference Performance Engineer (LLM)

AMD • Helsinki

Hybrid
EUR 140,000 - 200,000
Senior AI Performance Engineer - LLM Inference (vLLM)
Senior AI Performance Engineer - LLM Inference (vLLM)

Advanced Micro Devices • Helsinki

On-site
EUR 120,000 - 180,000
AMD benefits
Senior LLM Performance & Inference Lead
Senior LLM Performance & Inference Lead

AMD • Helsinki

On-site
EUR 90,000 - 130,000
Senior LLM Inference Performance Engineer
Senior LLM Inference Performance Engineer

AMD • Helsinki

On-site
EUR 90,000 - 130,000
Senior AI Inference Performance Engineer — SGLang
Senior AI Inference Performance Engineer — SGLang

Advanced Micro Devices • Helsinki

On-site
EUR 150,000 - 190,000
Senior LLM Inference Performance Engineer
Senior LLM Inference Performance Engineer

AMD • Helsinki

On-site
EUR 90,000 - 150,000
Principal AI Performance Engineer - LLM Inference (SGLang)
Principal AI Performance Engineer - LLM Inference (SGLang)

AMD • Helsinki

On-site
EUR 90,000 - 150,000
Principal AI Performance Engineer - LLM Inference (SGLang)
Principal AI Performance Engineer - LLM Inference (SGLang)

Advanced Micro Devices • Helsinki

On-site
EUR 150,000 - 190,000
Edge AI Engineer for Real-Time Embedded Mobility
Edge AI Engineer for Real-Time Embedded Mobility

XpertDirect • Helsinki

On-site
EUR 90,000 - 130,000