Stand out for this role — generate a tailored resume and cover letter in about a minute.
XG Tech PTE.LTD. in the Philippines is seeking a Large Model Quantization Algorithm Engineer to develop and optimize quantization and model compression for LLMs, VLMs, and video models.
You will improve accuracy, reduce memory use, and boost on-device inference across NPUs, GPUs, and CPUs, collaborating with compiler and hardware teams for production deployment. The role emphasizes PTQ/QAT methods, cross-hardware adaptation, and building automation tools for quantization workflows, with a focus
Founded in 2022, XG Tech is driving the future of smart vehicles. Its mission is to empower the digital transformation of automobiles, moving from distributed computing to a centralized, cross-domain platform.
XG Tech focuses on the intelligent cockpit—the next frontier of differentiation—while seamlessly integrating advanced driving systems. By reimagining cars as mobile living spaces, XG Tech aligns with the evolving trend of vehicles becoming the “third living space.”
As a Large Model Quantization Algorithm Engineer, you will develop quantization and model compression algorithms for LLMs, VLMs, and video generation models. You will optimize model accuracy, memory efficiency, and inference performance across NPUs, GPUs, and CPUs, bridging the gap between model algorithms and on-device deployment. You will work closely with algorithm, compiler, and hardware teams to bring efficient AI inference technologies into production.