Member of Technical Staff, Inference

INFERACT SINGAPORE PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Inferact is seeking an inference runtime engineer to push the boundaries of LLM and diffusion model serving. You will optimize how models execute across diverse hardware and architectures, and work at the core of vLLM to accelerate AI inference.

The role requires deep knowledge of transformer models, strong Python and PyTorch skills, and experience with LLM inference systems. You will read papers, implement techniques, and contribute robust, maintainable code to complex ML systems.

Qualifications

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Deep understanding of transformer architectures and their variants.
  • Strong programming skills in Python with experience in PyTorch internals.
  • Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).
  • Ability to read and implement model architectures and inference techniques from research papers.
  • Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases.

Responsibilities

  • Push the boundaries of what's possible in LLM and diffusion model serving.
  • Optimize how models execute across diverse hardware and architectures.
  • Work at the core of vLLM to accelerate AI inference.

Skills

Python programming
PyTorch internals
Transformer architectures
LLM inference systems
Debug ML codebases

Education

Bachelor's degree in CS/Engineering

Tools

TensorRT-LLM
SGLang
TGI

Job description

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

About the Role

We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference.

Skills and Qualifications

Minimum qualifications:

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Deep understanding of transformer architectures and their variants.
  • Strong programming skills in Python with experience in PyTorch internals.
  • Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).
  • Ability to read and implement model architectures and inference techniques from research papers.
  • Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases.

Preferred qualifications:

  • Deep understanding of KV-cache memory management, prefix caching, and hybrid model serving.
  • Familiarity with RL frameworks and algorithms for LLMs.
  • Experience with multimodal inference (audio/image/video/text).
  • Contributions to open-source ML or system infrastructure projects.

Bonus points if you have:

  • Implemented core features in vLLM or other inference engine projects.
  • Contributed to vLLM integrations (verl, OpenRLHF, Unsloth, LlamaFactory, etc).
  • Written widely-shared technical blogs or side projects on vLLM or LLM inference.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

LLM Inference Runtime Engineer
LLM Inference Runtime Engineer

INFERACT SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Embedded LLM Systems Engineer
Embedded LLM Systems Engineer

Desay SV • Singapore

On-site
SGD 120,000 - 180,000
Member of Technical Staff, TPU Performance Engineering
Member of Technical Staff, TPU Performance Engineering

Inferact • Singapore

On-site
SGD 200,000 - 400,000
Medical, dental, and vision coverage
Equity options
Member of Technical Staff, AMD GPU Performance Engineering
Member of Technical Staff, AMD GPU Performance Engineering

Inferact • Singapore

On-site
SGD 200,000 - 400,000
Medical coverage
Dental coverage
Vision coverage
+1
Senior Inference Runtime Engineer
Senior Inference Runtime Engineer

Bitdeer Group • Singapore

On-site
SGD 150,000 - 210,000
Senior AI Engineer
Senior AI Engineer

Patsnap • Singapore

On-site
SGD 80,000 - 120,000
Lead Engineer/ Engineer, MLOps / SRE (Developer Experience), xCloud
Lead Engineer/ Engineer, MLOps / SRE (Developer Experience), xCloud

HTX (Home Team Science & Technology Agency) • Singapore

On-site
SGD 120,000 - 180,000
Senior LLM Inference Performance Engineer
Senior LLM Inference Performance Engineer

Bitdeer Group • Singapore

On-site
SGD 150,000 - 210,000
Senior LLM Inference Performance Engineer
Senior LLM Inference Performance Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
Welfare benefits
Training and mentoring
LLM Engineer
LLM Engineer

TechKnowledgey Pte Ltd • Singapore

On-site
SGD 150,000 - 190,000