Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.
Inferact is seeking an AMD GPU performance engineer to advance vLLM as a premier inference engine on AMD accelerators. You will build and optimize AMD GPU backends, kernels, and benchmarking infrastructure using ROCm, HIP, Triton, CK, and AITER.
You will work at the boundary of inference systems, kernels, compilers, and hardware, improving attention, GEMM, sampling, KV cache, and other communication-heavy paths to deliver fast, scalable inference and maintainable backend.
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.
We're looking for an AMD GPU performance engineer to make vLLM a first-class inference engine across the AMD accelerator ecosystem. You'll build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure using ROCm, HIP, Triton, CK, AITER, and related tooling so vLLM can deliver frontier inference performance on AMD GPUs.
You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving performance-critical paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations. Your work will help make AMD GPU support in vLLM usable, fast, benchmarked, and maintainable.
Fresh graduates are welcome to apply. Minimum years of work experience: 0.
Compensation: Monthly salary of S$15,000 to S$30,000, depending on background, skills, and experience, plus equity.
Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage.