Large Model Quantization Algorithm Engineer

XG TECH PTE.LTD.

Santo Niño 1st

On-site

PHP 1,674,000 - 2,344,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

XG Tech PTE.LTD. in the Philippines is seeking a Large Model Quantization Algorithm Engineer to develop and optimize quantization and model compression for LLMs, VLMs, and video models.

You will improve accuracy, reduce memory use, and boost on-device inference across NPUs, GPUs, and CPUs, collaborating with compiler and hardware teams for production deployment. The role emphasizes PTQ/QAT methods, cross-hardware adaptation, and building automation tools for quantization workflows, with a focus

Qualifications

  • Bachelor’s degree or above in a quantitative field.
  • Familiarity with LLM/VLM algorithms and deployment optimization techniques, including model quantization, sparsity/pruning, and inference acceleration frameworks.
  • Prior hands-on experience with PyTorch Quantization-Aware Training (QAT) development is advantageous.
  • Strong proficiency in Python and C++.

Responsibilities

  • Develop and optimize quantization algorithms for LLMs, VLMs, and video generation models, covering PTQ, QAT, and related model compression techniques.
  • Design and evaluate quantization schemes to balance model accuracy, inference performance, and memory efficiency.
  • Perform quantization calibration and error analysis, identifying sources of accuracy degradation and driving optimization solutions.
  • Optimize edge and on-device inference, including Prefill/Decode acceleration, KV Cache management, operator fusion, weight compression, and memory optimization.
  • Adapt and deploy models across heterogeneous hardware, including NPU and CPU platforms, working closely with compiler teams on model conversion, engine compilation, and performance tuning.
  • Develop quantization and model optimization toolkits, including automated quantization workflows, accuracy evaluation, and visualization/debugging tools.
  • Collaborate with model and architecture teams to develop quantization-friendly model architectures, training strategies, and inference optimization techniques.
  • Track and evaluate emerging research in model quantization, compression, sparsity, and efficient inference, and drive relevant techniques into production.

Skills

Python
C++
PyTorch QAT
Quantization
Edge deployment

Education

Bachelor’s degree or above in Electronics Engineering, Computer Science, Automation, Operations Research, Statistics, Mathematics

Job description

About Company

Founded in 2022, XG Tech is driving the future of smart vehicles. Its mission is to empower the digital transformation of automobiles, moving from distributed computing to a centralized, cross-domain platform.

XG Tech focuses on the intelligent cockpit—the next frontier of differentiation—while seamlessly integrating advanced driving systems. By reimagining cars as mobile living spaces, XG Tech aligns with the evolving trend of vehicles becoming the “third living space.”

Role Summary

As a Large Model Quantization Algorithm Engineer, you will develop quantization and model compression algorithms for LLMs, VLMs, and video generation models. You will optimize model accuracy, memory efficiency, and inference performance across NPUs, GPUs, and CPUs, bridging the gap between model algorithms and on-device deployment. You will work closely with algorithm, compiler, and hardware teams to bring efficient AI inference technologies into production.

Key Responsibilities
  • Develop and optimize quantization algorithms for LLMs, VLMs, and video generation models, covering PTQ, QAT, and related model compression techniques.
  • Design and evaluate quantization schemes to balance model accuracy, inference performance, and memory efficiency.
  • Perform quantization calibration and error analysis, identifying sources of accuracy degradation and driving optimization solutions.
  • Optimize edge and on-device inference, including Prefill/Decode acceleration, KV Cache management, operator fusion, weight compression, and memory optimization.
  • Adapt and deploy models across heterogeneous hardware, including NPU and CPU platforms, working closely with compiler teams on model conversion, engine compilation, and performance tuning.
  • Develop quantization and model optimization toolkits, including automated quantization workflows, accuracy evaluation, and visualization/debugging tools.
  • Collaborate with model and architecture teams to develop quantization-friendly model architectures, training strategies, and inference optimization techniques.
  • Track and evaluate emerging research in model quantization, compression, sparsity, and efficient inference, and drive relevant techniques into production.
How will you stand out
  • Bachelor’s degree or above in Electronic Engineering, Computer Science, Automation, Operations Research, Statistics, Mathematics, or a related quantitative field.
  • Familiarity with LLM/VLM algorithms and deployment optimization techniques, including model quantization, sparsity/pruning, and inference acceleration frameworks.
  • Prior hands-on experience with PyTorch Quantization-Aware Training (QAT) development is advantageous.
  • Strong proficiency in Python and C++.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

On-Device AI Quantization Engineer for LLMs & VLMs
On-Device AI Quantization Engineer for LLMs & VLMs

XG TECH PTE.LTD. • Santo Niño 1st

On-site
PHP 1,674,000 - 2,344,000
AI Research Engineer (Model Compression & Quantization)
AI Research Engineer (Model Compression & Quantization)

Tether.io • España

On-site
PHP 4,914,004 - 7,371,007
Image Algorithm Engineer
Image Algorithm Engineer

XG TECH PTE.LTD. • Santo Niño 1st

On-site
PHP 900,000 - 1,200,000
AI Research Engineer, Model Compression & Quantization
AI Research Engineer, Model Compression & Quantization

Tether.io • España

On-site
PHP 4,914,004 - 7,371,007
Large Language Model Algorithm Engineer
Large Language Model Algorithm Engineer

Binance • Hinoba-an

On-site
PHP 1,200,000 - 2,400,000
Research Intern — Controllable Video Generation & Diffusion Acceleration
Research Intern — Controllable Video Generation & Diffusion Acceleration

XG TECH PTE.LTD. • Santo Niño 1st

On-site
PHP 123,000 - 201,000
AI Application Engineer (Part-Time)
AI Application Engineer (Part-Time)

Workana • Mexico

On-site
PHP 3,743,000 - 7,486,000
Member of Technical Staff, LLM Infrastructure
Member of Technical Staff, LLM Infrastructure

Fireworks • San Mateo

On-site
PHP 1,800,000 - 3,000,000
LLM Engineer
LLM Engineer

KDCI Outsourcing • Pasig

On-site
PHP 1,000,000 - 2,000,000
AI Engineer – LLM Algorithm Engineer (Agentic Commerce)
AI Engineer – LLM Algorithm Engineer (Agentic Commerce)

Hammerjack Pty Ltd • Philippines

On-site
PHP 2,000,000 - 4,500,000