Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Rakuten Asia Pte. Ltd. in Singapore is seeking an LLM Inference Optimization Engineer to maximize the performance, efficiency, and scalability of large-scale inference workloads on GPU clusters. You will optimize engines and GPU kernels to ensure peak model serving efficiency.
The role requires deep expertise in CUDA/Triton, inference internals, and experience with at least one serving engine. Collaboration with global teams across Rakuten will be essential.
Rakuten Asia Pte. Ltd. in Singapore is seeking an LLM Inference Optimization Engineer to maximize the performance, efficiency, and scalability of large-scale inference workloads on GPU clusters. You will optimize engines and GPU kernels to ensure peak model serving efficiency.
The role requires deep expertise in CUDA/Triton, inference internals, and experience with at least one serving engine. Collaboration with global teams across Rakuten will be essential.