On-Device AI Quantization Engineer for LLMs & VLMs

XG TECH PTE.LTD.

Santo Niño 1st

On-site

PHP 1,674,000 - 2,344,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

XG Tech PTE.LTD. in the Philippines is seeking a Large Model Quantization Algorithm Engineer to develop and optimize quantization and model compression for LLMs, VLMs, and video models.

You will improve accuracy, reduce memory use, and boost on-device inference across NPUs, GPUs, and CPUs, collaborating with compiler and hardware teams for production deployment. The role emphasizes PTQ/QAT methods, cross-hardware adaptation, and building automation tools for quantization workflows, with a focus

Qualifications

  • Bachelor’s degree or above in a quantitative field.
  • Familiarity with LLM/VLM algorithms and deployment optimization techniques, including model quantization, sparsity/pruning, and inference acceleration frameworks.
  • Prior hands-on experience with PyTorch Quantization-Aware Training (QAT) development is advantageous.
  • Strong proficiency in Python and C++.

Responsibilities

  • Develop and optimize quantization algorithms for LLMs, VLMs, and video generation models, covering PTQ, QAT, and related model compression techniques.
  • Design and evaluate quantization schemes to balance model accuracy, inference performance, and memory efficiency.
  • Perform quantization calibration and error analysis, identifying sources of accuracy degradation and driving optimization solutions.
  • Optimize edge and on-device inference, including Prefill/Decode acceleration, KV Cache management, operator fusion, weight compression, and memory optimization.
  • Adapt and deploy models across heterogeneous hardware, including NPU and CPU platforms, working closely with compiler teams on model conversion, engine compilation, and performance tuning.
  • Develop quantization and model optimization toolkits, including automated quantization workflows, accuracy evaluation, and visualization/debugging tools.
  • Collaborate with model and architecture teams to develop quantization-friendly model architectures, training strategies, and inference optimization techniques.
  • Track and evaluate emerging research in model quantization, compression, sparsity, and efficient inference, and drive relevant techniques into production.

Skills

Python
C++
PyTorch QAT
Quantization
Edge deployment

Education

Bachelor’s degree or above in Electronics Engineering, Computer Science, Automation, Operations Research, Statistics, Mathematics

Job description

XG Tech PTE.LTD. in the Philippines is seeking a Large Model Quantization Algorithm Engineer to develop and optimize quantization and model compression for LLMs, VLMs, and video models.

You will improve accuracy, reduce memory use, and boost on-device inference across NPUs, GPUs, and CPUs, collaborating with compiler and hardware teams for production deployment. The role emphasizes PTQ/QAT methods, cross-hardware adaptation, and building automation tools for quantization workflows, with a focus

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Large Model Quantization Algorithm Engineer
Large Model Quantization Algorithm Engineer

XG TECH PTE.LTD. • Santo Niño 1st

On-site
PHP 1,674,000 - 2,344,000
AI Research Engineer (Model Compression & Quantization)
AI Research Engineer (Model Compression & Quantization)

Tether.io • España

On-site
PHP 4,914,004 - 7,371,007
AI Research Engineer, Model Compression & Quantization
AI Research Engineer, Model Compression & Quantization

Tether.io • España

On-site
PHP 4,914,004 - 7,371,007
AI Engineer: Build Production ML & LLM Solutions
AI Engineer: Build Production ML & LLM Solutions

Cctech • Hinoba-an

On-site
PHP 900,000 - 1,500,000
Cutting-edge AI projects
Growth opportunities
ML Engineer: NLP, TensorFlow & LLMs (Hands-On)
ML Engineer: NLP, TensorFlow & LLMs (Hands-On)

TymblHub • Hinoba-an

On-site
PHP 400,000 - 640,000
Senior GenAI Engineer: Build Scalable LLM Solutions
Senior GenAI Engineer: Build Scalable LLM Solutions

TymblHub • Hinoba-an

On-site
PHP 1,674,000 - 3,571,000
Production AI/ML Engineer: Build & Deploy Smart Models
Production AI/ML Engineer: Build & Deploy Smart Models

Quik Hire Staffing • Philippines

On-site
PHP 900,000 - 1,400,000
Production GenAI Engineer for Enterprise LLMs
Production GenAI Engineer for Enterprise LLMs

Onebyzero • Philippines

Hybrid
PHP 1,200,000 - 1,800,000
Remote LLM Fine-Tuning Engineer for Production
Remote LLM Fine-Tuning Engineer for Production

Workana • Mexico

Hybrid
PHP 3,743,000 - 7,486,000
Senior Quality Engineer - AI & Data Testing Lead
Senior Quality Engineer - AI & Data Testing Lead

TASQ • Makati

On-site
PHP 600,000 - 900,000