Embedded LLM Systems Engineer

Desay SV

Singapore

On-site

SGD 120,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Desay SV is seeking an Embedded LLM Systems Engineer to design and optimise on-device LLM inference for embedded, mobile and edge devices. You will work on inference engines, operator development and graph optimisation across multiple backends.

The role requires hands-on experience with modern C++, Python, and model-quantisation techniques, plus familiarity with CUDA, MediaPipe and related technologies for efficient deployment.

Qualifications

  • Three years of experience in on-device inference, AI infrastructure, embedded systems engineering, or related area.
  • Bachelor’s degree or above in CS/EE/Math or equivalent practical experience.
  • Proficiency in English to read technical docs and discuss with teams.

Responsibilities

  • Design, develop and optimise LLM inference engines for embedded, mobile and edge devices.
  • Work on operator development, graph optimisation, memory management and multi-backend adaptation.
  • Develop solutions using frameworks such as llama.cpp, TensorRT-LLM, MNN, ONNX Runtime or comparable technologies.
  • Research and apply quantisation techniques such as INT4, INT8 and FP16.
  • Work with technologies such as NEON/SVE, Vulkan Compute, OpenCL or comparable platforms.
  • Conduct training-inference consistency validation and support deployment across cloud and edge environments.
  • Translate emerging AI capabilities into embedded product value.

Skills

On-device inference
Modern C++
Python
English proficiency

Education

Bachelor's degree in Computer Science
Bachelor's degree in Electrical/Electronic Engineering
Bachelor's degree in Mathematics

Tools

llama.cpp
TensorRT-LLM
MNN
ONNX Runtime
CUDA
MediaPipe

Job description

We are seeking an Embedded LLM Systems Engineer to design, develop and optimise LLM inference solutions for embedded, mobile and edge devices. The role covers inference-engine development, model compression, heterogeneous hardware optimisation and efficient model deployment.

Duties/ Responsibilities:

On-Device Inference Engine Development

  • Design, develop and optimise LLM inference engines for embedded, mobile and edge devices.
  • Work on operator development, graph optimisation, memory management and multi-backend adaptation.
  • Develop solutions using frameworks such as llama.cpp, TensorRT-LLM, MNN, ONNX Runtime or comparable technologies.
  • Research and apply quantisation techniques such as INT4, INT8 and FP16.
  • Work with relevant technologies such as NEON/SVE, Vulkan Compute, OpenCL or comparable platforms.
  • Conduct training-inference consistency validation and support efficient deployment across cloud and edge environments.
  • Evaluate practical ways to translate emerging AI capabilities into embedded product applications.
Basic Requirements:
  • At least three years of relevant experience in on-device inference, AI infrastructure, embedded systems engineering or a related area.
  • Bachelor’s degree or above in Computer Science, Electrical/Electronic Engineering, Mathematics or a related discipline, or equivalent practical experience.
  • Proficiency in written and spoken English sufficient to read technical documentation and research papers, participate in technical discussions and prepare clear engineering documentation.
  • Strong proficiency in modern C++, with a good understanding of memory models, concurrency and low-level performance optimisation.
  • Proficiency in Python for model conversion, evaluation, automation scripts and training-related tooling.
  • Experience with CUDA, MediaPipe or related technologies is advantageous.
Preferred Qualifications

Any of the following would be advantageous:

  • Contributions to established open-source inference projects such as llama.cpp, vLLM, TensorRT-LLM, MLC-LLM or MNN.
  • Publications in recognised conferences or journals on efficient inference, model compression or on-device deployment.
  • Recognition in competitions such as ACM-ICPC, NOI, Kaggle or on-device AI challenges.
  • Experience with prompt engineering, Retrieval-Augmented Generation or AI agent frameworks such as LangChain or LlamaIndex.
  • Hands-on experience deploying and optimising inference frameworks such as vLLM, TGI, llama.cpp, TensorRT-LLM or MLC-LLM.
  • Knowledge of model alignment and fine-tuning techniques, including RLHF, SFT and DPO.
  • Practical experience fine-tuning or evaluating models using authorised, organisation-owned or appropriately licensed datasets.
  • Strong interest in emerging LLM technologies and the ability to translate new capabilities into practical product value.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Edge LLM Engineer: On-Device Inference
Edge LLM Engineer: On-Device Inference

Desay SV • Singapore

On-site
SGD 120,000 - 180,000
Member of Technical Staff, Inference
Member of Technical Staff, Inference

INFERACT SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
LLM Inference Runtime Engineer
LLM Inference Runtime Engineer

INFERACT SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Sr. AI/ML Engineer
Sr. AI/ML Engineer

GECO Asia Pte Ltd • Singapore

Hybrid
SGD 150,000 - 210,000
Senior Software Engineer
Senior Software Engineer

Rakuten Asia Pte Ltd • Singapore

On-site
SGD 140,000 - 200,000
AI Application Engineer
AI Application Engineer

AZTECH TECHNOLOGIES PTE LTD • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Full-Stack Engineer — LLMs & Cloud
Senior AI Full-Stack Engineer — LLMs & Cloud

KEYSIGHT TECHNOLOGIES SINGAPORE (SALES) PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
LLM Engineer
LLM Engineer

TechKnowledgey Pte Ltd • Singapore

On-site
SGD 150,000 - 190,000
LLM Engineer
LLM Engineer

TECHKNOWLEDGEY PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Engineer
Senior AI Engineer

Patsnap • Singapore

On-site
SGD 80,000 - 120,000