Senior LLM Inference Optimization Engineer

Rakuten Asia Pte Ltd

Singapore

Hybrid

SGD 120,000 - 190,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Rakuten Asia Pte. Ltd. in Singapore is seeking a Senior Software Engineer focusing on LLM inference optimization.

You will profile, optimize, and extend inference engines to improve throughput, reduce latency, and maximize GPU utilization on NVIDIA GPUs. The role requires deep CUDA/Triton programming, experience with vLLM, TensorRT-LLM, or Triton Inference Server, plus strong Python and C++ software engineering.

Qualifications

  • 3+ years of hands-on experience optimizing deep learning inference on NVIDIA GPUs.
  • Strong knowledge of LLM inference internals: attention mechanisms, KV cache, batching strategies, quantization, and parallelism.
  • Working experience with at least one mainstream inference serving engine (vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or equivalent).
  • Proficiency in CUDA and/or Triton kernel programming, with a solid understanding of modern GPU architecture.
  • Strong software engineering fundamentals in Python and C++.
  • Bachelor's or higher degree in Computer Science, Engineering, or a related field.

Responsibilities

  • Profile, optimize, and extend modern LLM inference engines (e.g., vLLM, SGLang, TensorRT-LLM) to improve throughput, latency, and GPU utilization.
  • Design and implement inference-time optimizations such as quantization, KV-cache management, continuous batching, speculative decoding, and parallelism strategies (tensor / expert / pipeline).
  • Develop and tune GPU kernels (CUDA, Triton) for critical operators — attention, GEMM, MoE — targeting the latest NVIDIA architectures.
  • Build benchmarking and load-testing frameworks to measure end-to-end serving performance under realistic workloads and SLOs.
  • Partner with infrastructure teams to deploy and scale inference on Kubernetes-based GPU clusters, including autoscaling, fault tolerance, and observability.
  • Stay current with the fast-moving inference research landscape and bring cutting-edge techniques into production.

Skills

LLM inference optimization
CUDA
Python
C++
Triton
GPU kernels
quantization
KV cache
batching strategies
inference serving

Education

Bachelor's or higher in CS/Engineering

Tools

vLLM
SGLang
TensorRT-LLM
Triton Inference Server
CUDA

Job description

Rakuten Asia Pte. Ltd. in Singapore is seeking a Senior Software Engineer focusing on LLM inference optimization.

You will profile, optimize, and extend inference engines to improve throughput, reduce latency, and maximize GPU utilization on NVIDIA GPUs. The role requires deep CUDA/Triton programming, experience with vLLM, TensorRT-LLM, or Triton Inference Server, plus strong Python and C++ software engineering.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer
Senior Software Engineer

Rakuten Asia Pte Ltd • Singapore

Hybrid
SGD 120,000 - 190,000
Senior LLM Inference Performance Engineer
Senior LLM Inference Performance Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
Welfare benefits
Training and mentoring
LLM Platform Engineer - AI Inference & Ops
LLM Platform Engineer - AI Inference & Ops

Nanyang Technological University Singapore • Singapore

On-site
SGD 90,000 - 140,000
LLM Inference Runtime Engineer
LLM Inference Runtime Engineer

INFERACT SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Embedded LLM Systems Engineer
Embedded LLM Systems Engineer

Desay SV • Singapore

On-site
SGD 120,000 - 180,000
LLM-Driven AI Infrastructure Intern
LLM-Driven AI Infrastructure Intern

TikTok • Singapore

On-site
SGD 13,000 - 20,000
LLM Pre-Training Engineer — Scale & Optimize
LLM Pre-Training Engineer — Scale & Optimize

UMELIFE (SINGAPORE) PTE. LTD. • Singapore

On-site
SGD 180,000 - 260,000
AI/ML Engineer: LLM Training & Deployment
AI/ML Engineer: LLM Training & Deployment

K.P.P. PACKAGING PTE LTD • Singapore

On-site
SGD 80,000 - 140,000
Senior LLM Inference Performance Engineer
Senior LLM Inference Performance Engineer

Bitdeer Group • Singapore

On-site
SGD 150,000 - 210,000
TPU Inference Performance Engineer
TPU Inference Performance Engineer

Inferact • Singapore

On-site
SGD 200,000 - 400,000
Medical, dental, and vision coverage
Equity options