Stand out for this role — generate a tailored resume and cover letter in about a minute.
Inferact seeks an AMD GPU performance engineer to advance vLLM as a premier inference engine on AMD accelerators. You will build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure to deliver frontier inference performance on AMD GPUs.
You will work at the intersection of inference systems, kernels, compilers, and hardware architecture, focusing on attention, GEMM, sampling, KV cache, and communication-heavy paths to enhance speed, reliability, and
Inferact seeks an AMD GPU performance engineer to advance vLLM as a premier inference engine on AMD accelerators. You will build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure to deliver frontier inference performance on AMD GPUs.
You will work at the intersection of inference systems, kernels, compilers, and hardware architecture, focusing on attention, GEMM, sampling, KV cache, and communication-heavy paths to enhance speed, reliability, and