Staff AMD GPU Performance Engineer, Inference Acceleration

INFERACT SINGAPORE PTE. LTD.

Penarth

On-site

GBP 99,000 - 198,000

Full time

10 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Medical coverage
Yearly bonus

Job summary

Inferact is seeking an AMD GPU performance engineer to advance vLLM as a premier inference engine on AMD accelerators. You will build and optimize AMD GPU backends, kernels, and benchmarking infrastructure using ROCm, HIP, Triton, CK, and AITER.

You will work at the boundary of inference systems, kernels, compilers, and hardware, improving attention, GEMM, sampling, KV cache, and other communication-heavy paths to deliver fast, scalable inference and maintainable backend.

Qualifications

  • Bachelor's degree or equivalent in CS, engineering, systems, ML, or similar.
  • Hands-on experience optimizing AMD GPU workloads using ROCm, HIP, Triton, CK, AITER, or similar tools.
  • Deep understanding of AMD GPU execution, memory behavior, toolchains, kernel performance, and backend constraints.
  • Experience optimizing ML kernels or inference paths (attention, GEMM, sampling, KV cache, fused kernels, or heavy-communication paths).
  • Strong performance profiling with measurements, hardware counters, correctness tests, and reproducible benchmarks.

Responsibilities

  • Build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure using ROCm, HIP, Triton, CK, AITER, and related tooling.
  • Improve performance-critical paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations to enable fast, maintainable backend for vLLM.

Skills

ROCm
HIP
Triton
CK
AITER
Performance profiling
ML kernels

Education

Bachelor's degree or equivalent

Tools

vLLM
TensorRT-LLM
LLVM
MLIR

Job description

Inferact is seeking an AMD GPU performance engineer to advance vLLM as a premier inference engine on AMD accelerators. You will build and optimize AMD GPU backends, kernels, and benchmarking infrastructure using ROCm, HIP, Triton, CK, and AITER.

You will work at the boundary of inference systems, kernels, compilers, and hardware, improving attention, GEMM, sampling, KV cache, and other communication-heavy paths to deliver fast, scalable inference and maintainable backend.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, AMD GPU Performance Engineering
Member of Technical Staff, AMD GPU Performance Engineering

INFERACT SINGAPORE PTE. LTD. • Penarth

On-site
GBP 99,000 - 198,000
Equity
Medical coverage
Yearly bonus
Performance Kernel Engineer, GPU & ML Inference
Performance Kernel Engineer, GPU & ML Inference

INFERACT SINGAPORE PTE. LTD. • Penarth

On-site
GBP 99,000 - 198,000
Equity
Medical insurance
Dental insurance
+1
Member of Technical Staff, Kernel Engineering
Member of Technical Staff, Kernel Engineering

INFERACT SINGAPORE PTE. LTD. • Penarth

On-site
GBP 99,000 - 198,000
Equity
Medical insurance
Dental insurance
+1
GPU Performance Engineer: Scale Inference & Training
GPU Performance Engineer: Scale Inference & Training

Anthropic • York and North Yorkshire

On-site
GBP 90,000 - 140,000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1
Senior GPU HPC Infrastructure Engineer (Equity)
Senior GPU HPC Infrastructure Engineer (Equity)

INFERACT SINGAPORE PTE. LTD. • Penarth

On-site
GBP 99,000 - 198,000
Medical insurance
Dental insurance
Vision insurance
Member of Technical Staff, TPU Performance Engineering
Member of Technical Staff, TPU Performance Engineering

INFERACT SINGAPORE PTE. LTD. • Penarth

On-site
GBP 99,000 - 198,000
Equity
Medical coverage
Dental coverage
+1
TPU Performance Engineer - vLLM Inference Backends
TPU Performance Engineer - vLLM Inference Backends

INFERACT SINGAPORE PTE. LTD. • Penarth

On-site
GBP 99,000 - 198,000
Equity
Medical coverage
Dental coverage
+1
Senior AI Compute Kernels Engineer - Performance
Senior AI Compute Kernels Engineer - Performance

EngineersOfAI • Bristol

On-site
GBP 70,000 - 120,000
Member of Technical Staff, Performance and Scale
Member of Technical Staff, Performance and Scale

INFERACT SINGAPORE PTE. LTD. • Penarth

On-site
GBP 99,000 - 198,000
Medical coverage
Dental coverage
Vision coverage
+1
Performance Engineer (GPU)
Performance Engineer (GPU)

Anthropic • York and North Yorkshire

On-site
GBP 90,000 - 140,000
Comprehensive health insurance
Fertility benefits
22 weeks parental leave
+1