GPU-Accelerated LLM Inference Engineer

Rakuten Asia Pte Ltd

Singapore

On-site

SGD 140,000 - 200,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Rakuten Asia Pte. Ltd. in Singapore is seeking an LLM Inference Optimization Engineer to maximize the performance, efficiency, and scalability of large-scale inference workloads on GPU clusters. You will optimize engines and GPU kernels to ensure peak model serving efficiency.

The role requires deep expertise in CUDA/Triton, inference internals, and experience with at least one serving engine. Collaboration with global teams across Rakuten will be essential.

Qualifications

  • 3+ years of hands-on experience optimizing deep learning inference on NVIDIA GPUs, preferably for large language models.
  • Strong knowledge of LLM inference internals: attention mechanisms, KV cache, batching strategies, quantization, and parallelism.
  • Working experience with at least one mainstream inference serving engine (vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or equivalent).
  • Proficiency in CUDA and/or Triton kernel programming, with a solid understanding of modern GPU architecture.
  • Strong software engineering fundamentals in Python and C++.
  • Bachelor's or higher degree in Computer Science, Engineering, or a related field.

Responsibilities

  • Profile, optimize, and extend modern LLM inference engines (e.g., vLLM, SGLang, TensorRT-LLM) to improve throughput, latency, and GPU utilization.
  • Design and implement inference-time optimizations such as quantization, KV-cache management, continuous batching, speculative decoding, and parallelism strategies (tensor / expert / pipeline).
  • Develop and tune GPU kernels (CUDA, Triton) for critical operators — attention, GEMM, MoE — targeting the latest NVIDIA architectures.
  • Build benchmarking and load-testing frameworks to measure end-to-end serving performance under realistic workloads and SLOs.
  • Partner with infrastructure teams to deploy and scale inference on Kubernetes-based GPU clusters, including autoscaling, fault tolerance, and observability.
  • Stay current with the fast-moving inference research landscape and bring cutting-edge techniques into production.

Skills

CUDA
Python
C++
GPU kernels
Inference optimization
Inference engines

Education

Bachelor's degree in CS/Engineering

Tools

Triton
Nsight Systems
PyTorch
TensorRT
Kubernetes
DeepSpeed

Job description

Rakuten Asia Pte. Ltd. in Singapore is seeking an LLM Inference Optimization Engineer to maximize the performance, efficiency, and scalability of large-scale inference workloads on GPU clusters. You will optimize engines and GPU kernels to ensure peak model serving efficiency.

The role requires deep expertise in CUDA/Triton, inference internals, and experience with at least one serving engine. Collaboration with global teams across Rakuten will be essential.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Software Engineer
Senior Software Engineer

Rakuten Asia Pte Ltd • Singapore

On-site
SGD 140,000 - 200,000
AI Training & Inference Acceleration Architect
AI Training & Inference Acceleration Architect

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
EDB-IPP Project: Advancing GPU Optimization for Large Language Models
EDB-IPP Project: Advancing GPU Optimization for Large Language Models

Rakuten Kobo Inc. • Singapore

On-site
SGD 42,000 - 54,000
ML Systems Engineer: TPU Training & Inference Optimizer
ML Systems Engineer: TPU Training & Inference Optimizer

GOOGLE ASIA PACIFIC PTE. LTD. • Singapore

On-site
SGD 180,000 - 260,000
AI Computing Architect for LLM & Distributed AI
AI Computing Architect for LLM & Distributed AI

PERSOL SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Staff Research Software Engineer: AI Inference & LLMs
Staff Research Software Engineer: AI Inference & LLMs

Google Inc. • Singapore

On-site
SGD 180,000 - 240,000
Staff ML Engineer: TPU Training & Inference Optimization
Staff ML Engineer: TPU Training & Inference Optimization

Google • Singapore

On-site
SGD 180,000 - 240,000
TPU Inference Performance Engineer
TPU Inference Performance Engineer

Inferact • Singapore

On-site
SGD 200,000 - 400,000
Medical, dental, and vision coverage
Equity options
Senior LLM Inference Performance Engineer
Senior LLM Inference Performance Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
Welfare benefits
Training and mentoring
LLM Platform Engineer - AI Inference & Ops
LLM Platform Engineer - AI Inference & Ops

Nanyang Technological University Singapore • Singapore

On-site
SGD 90,000 - 140,000