A complete application in a minute — tailored resume and cover letter, ready to send.
Rakuten Asia Pte. Ltd. in Singapore is seeking a Senior Software Engineer focusing on LLM inference optimization.
You will profile, optimize, and extend inference engines to improve throughput, reduce latency, and maximize GPU utilization on NVIDIA GPUs. The role requires deep CUDA/Triton programming, experience with vLLM, TensorRT-LLM, or Triton Inference Server, plus strong Python and C++ software engineering.
Rakuten Asia Pte. Ltd. in Singapore is seeking a Senior Software Engineer focusing on LLM inference optimization.
You will profile, optimize, and extend inference engines to improve throughput, reduce latency, and maximize GPU utilization on NVIDIA GPUs. The role requires deep CUDA/Triton programming, experience with vLLM, TensorRT-LLM, or Triton Inference Server, plus strong Python and C++ software engineering.