LLM Inference Runtime Engineer

INFERACT SINGAPORE PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Inferact is seeking an inference runtime engineer to push the boundaries of LLM and diffusion model serving. You will optimize how models execute across diverse hardware and architectures, and work at the core of vLLM to accelerate AI inference.

The role requires deep knowledge of transformer models, strong Python and PyTorch skills, and experience with LLM inference systems. You will read papers, implement techniques, and contribute robust, maintainable code to complex ML systems.

Qualifications

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Deep understanding of transformer architectures and their variants.
  • Strong programming skills in Python with experience in PyTorch internals.
  • Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).
  • Ability to read and implement model architectures and inference techniques from research papers.
  • Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases.

Responsibilities

  • Push the boundaries of what's possible in LLM and diffusion model serving.
  • Optimize how models execute across diverse hardware and architectures.
  • Work at the core of vLLM to accelerate AI inference.

Skills

Python programming
PyTorch internals
Transformer architectures
LLM inference systems
Debug ML codebases

Education

Bachelor's degree in CS/Engineering

Tools

TensorRT-LLM
SGLang
TGI

Job description

Inferact is seeking an inference runtime engineer to push the boundaries of LLM and diffusion model serving. You will optimize how models execute across diverse hardware and architectures, and work at the core of vLLM to accelerate AI inference.

The role requires deep knowledge of transformer models, strong Python and PyTorch skills, and experience with LLM inference systems. You will read papers, implement techniques, and contribute robust, maintainable code to complex ML systems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Inference
Member of Technical Staff, Inference

INFERACT SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Embedded LLM Systems Engineer
Embedded LLM Systems Engineer

Desay SV • Singapore

On-site
SGD 120,000 - 180,000
Senior LLM Inference Performance Engineer
Senior LLM Inference Performance Engineer

Bitdeer Group • Singapore

On-site
SGD 150,000 - 210,000
Senior LLM Inference Performance Engineer
Senior LLM Inference Performance Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
Welfare benefits
Training and mentoring
Edge LLM Engineer: On-Device Inference
Edge LLM Engineer: On-Device Inference

Desay SV • Singapore

On-site
SGD 120,000 - 180,000
Member of Technical Staff, TPU Performance Engineering
Member of Technical Staff, TPU Performance Engineering

Inferact • Singapore

On-site
SGD 200,000 - 400,000
Medical, dental, and vision coverage
Equity options
Senior Software Engineer
Senior Software Engineer

Rakuten Asia Pte Ltd • Singapore

On-site
SGD 140,000 - 200,000
LLM-Driven AI Infrastructure Intern
LLM-Driven AI Infrastructure Intern

TikTok • Singapore

On-site
SGD 13,000 - 20,000
GPU-Accelerated LLM Inference Engineer
GPU-Accelerated LLM Inference Engineer

Rakuten Asia Pte Ltd • Singapore

On-site
SGD 140,000 - 200,000
LLM Engineer
LLM Engineer

TechKnowledgey Pte Ltd • Singapore

On-site
SGD 150,000 - 190,000