An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Inferact is hiring an AMD GPU performance engineer in Singapore to enhance vLLM for AMD accelerators. You'll develop and optimize AMD GPU backends using tools like ROCm and Triton to ensure top-tier performance.
Ideal candidates have a Bachelor's degree in a related field and hands-on experience with AMD GPU workloads. The role offers a competitive salary range of SGD 200,000 to 400,000 annually, along with comprehensive benefits including medical, dental, and vision coverage.
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.
We're looking for an AMD GPU performance engineer to make vLLM a first-class inference engine across the AMD accelerator ecosystem. You'll build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure using ROCm, HIP, Triton, CK, AITER, and related tooling so vLLM can deliver frontier inference performance on AMD GPUs.
You’ll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving performance‑critical paths such as attention, GEMM, sampling, KV cache, and communication‑heavy operations. Your work will help make AMD GPU support in vLLM usable, fast, benchmarked, and maintainable.
Minimum qualifications:
Preferred qualifications:
Bonus points if you have: