Senior AI Inference Performance Engineer (LLM)

AMD

Helsinki

Hybrid

EUR 140,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

AMD is seeking a senior engineer to push AI inference performance on GPUs using vLLM as the primary serving framework. You will own hard performance problems across the stack, profiling, diagnosing, and optimizing models on customer-engaged deployments, delivering measurable uplifts and reusable methodologies.

You will work end-to-end across profiling, kernel and system-level optimizations, contributing to shared performance methodologies and accelerating daily work with AI agents.

Qualifications

  • 5+ years of software development experience in GPU computing, AI systems, or high-performance computing.
  • Hands-on experience with vLLM internals (V1 engine, scheduler, PagedAttention/KV cache manager).
  • Strong profiling and bottleneck-diagnosis background across GPU kernels and system software.
  • Ability to translate profiling data into concrete optimizations and improvements.

Responsibilities

  • Drive performance optimization end-to-end on vLLM across leading models and configurations.
  • Profile, diagnose, and resolve cross-stack performance bottlenecks from GPU kernels to vLLM scheduler.
  • Diagnose kernel-level issues and translate findings into actionable optimizations.
  • Present findings and deliver measurable performance uplifts for customer engagements.
  • Integrate and optimize custom kernels within vLLM and related open-source frameworks.

Skills

GPU computing
AI systems
End-to-end profiling
Python
C++
Kernel optimization
RDMA
Multi-node

Education

Master's in CS/CE

Tools

HIP
CUDA
Triton
CK
Gluon
SGLang
PyTorch

Job description

AMD is seeking a senior engineer to push AI inference performance on GPUs using vLLM as the primary serving framework. You will own hard performance problems across the stack, profiling, diagnosing, and optimizing models on customer-engaged deployments, delivering measurable uplifts and reusable methodologies.

You will work end-to-end across profiling, kernel and system-level optimizations, contributing to shared performance methodologies and accelerating daily work with AI agents.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Inference Performance Engineer - vLLM on GPUs
Senior AI Inference Performance Engineer - vLLM on GPUs

Advanced Micro Devices • Helsinki

On-site
EUR 120,000 - 180,000
AMD benefits
Senior AI Performance Engineer - LLM Inference (vLLM)
Senior AI Performance Engineer - LLM Inference (vLLM)

AMD • Helsinki

Hybrid
EUR 140,000 - 200,000
Senior AI Performance Engineer - LLM Inference (vLLM)
Senior AI Performance Engineer - LLM Inference (vLLM)

Advanced Micro Devices • Helsinki

On-site
EUR 120,000 - 180,000
AMD benefits
Senior Technical Program Manager (TPM) – GenAI
Senior Technical Program Manager (TPM) – GenAI

Advanced Micro Devices • Helsinki

On-site
EUR 90,000 - 130,000
Senior GenAI TPM: Lead Forward-Deployed AI Programs
Senior GenAI TPM: Lead Forward-Deployed AI Programs

AMD • Helsinki

Hybrid
EUR 110,000 - 170,000
AMD benefits
Senior Technical Program Manager (TPM) – GenAI
Senior Technical Program Manager (TPM) – GenAI

AMD • Helsinki

Hybrid
EUR 110,000 - 170,000
AMD benefits
Senior Agentic System & Application Engineer
Senior Agentic System & Application Engineer

Advanced Micro Devices • Helsinki

Hybrid
EUR 90,000 - 130,000
Senior AI Agent Engineer — Production-Grade LLM Apps
Senior AI Agent Engineer — Production-Grade LLM Apps

Digital Workforce Services Plc • Helsinki

On-site
EUR 90,000 - 120,000
Great benefits package
Learning & development
Senior Technical Program Manager (TPM) – GenAI
Senior Technical Program Manager (TPM) – GenAI

AMD • Finland

Hybrid
EUR 110,000 - 170,000
Senior Agentic AI Systems Engineer for HPC & Enterprise
Senior Agentic AI Systems Engineer for HPC & Enterprise

AMD • Helsinki

Hybrid
EUR 90,000 - 140,000