AI Inference & Video Compression Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd.

Singapore

On-site

SGD 120,000 - 180,000

Full time

28 hours ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance

Job summary

Huawei is seeking an AI Inference & Compression Engineer in Singapore to advance LLM inference acceleration, develop advanced video codecs, and optimize AI-based media coding. The role includes deploying optimized models and integrating compression techniques into deployment frameworks such as vLLM.

The candidate will work on performance and quality evaluation, including PSNR, VMAF, perplexity, latency, and throughput, ensuring robust, scalable AI/video systems.

Qualifications

  • Master’s or PhD in Computer Science, Electronics, Mathematics, or related fields (PhD preferred).
  • Solid understanding of video coding fundamentals incl. prediction, transform coding, quantization, entropy coding (H.265/HEVC, AV1, H.266/VVC).
  • Strong understanding of Transformer architectures and attention mechanisms, memory bandwidth constraints in AI inference.
  • Proficiency in Python and C/C++, with hands-on experience in PyTorch, TensorFlow, etc.

Responsibilities

  • LLM Inference Acceleration: research and develop compression algorithms to accelerate LLM serving, KV cache optimization, decoding bottlenecks.
  • Classical Codec Development: improve RD performance, optimize entropy coding, enhance quantization for real-world apps.
  • AI-Based Media Coding: develop AI-based loop filters, optical flow, and rate control.
  • Model Deployment & Fusion: bridge AI research and production, optimize models for inference and integrate compression into frameworks like vLLM.
  • Performance & Quality Evaluation: assess PSNR, VMAF, perplexity, latency, and throughput for AI/video systems.

Skills

Python
C/C++
PyTorch
TensorFlow
Transformer
Memory bandwidth
H265/HEVC
AV1
H266/VVC

Education

Master's or PhD in CS/EE/Math

Tools

vLLM

Job description

Huawei is seeking an AI Inference & Compression Engineer in Singapore to advance LLM inference acceleration, develop advanced video codecs, and optimize AI-based media coding. The role includes deploying optimized models and integrating compression techniques into deployment frameworks such as vLLM.

The candidate will work on performance and quality evaluation, including PSNR, VMAF, perplexity, latency, and throughput, ensuring robust, scalable AI/video systems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Inference & Compression Engineer
AI Inference & Compression Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 120,000 - 180,000
Health insurance
AI Training & Inference Acceleration Architect
AI Training & Inference Acceleration Architect

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
AI Hardware Architect for Ultra-Efficient Inference
AI Hardware Architect for Ultra-Efficient Inference

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
AI Compute Architect: High-Performance Systems Lead
AI Compute Architect: High-Performance Systems Lead

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 260,000
AI Inference & Training Acceleration Architect
AI Inference & Training Acceleration Architect

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 280,000
AI Training/Inference Acceleration Algorithm Engineer
AI Training/Inference Acceleration Algorithm Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
GPU-Accelerated LLM Inference Engineer
GPU-Accelerated LLM Inference Engineer

Rakuten Asia Pte Ltd • Singapore

On-site
SGD 140,000 - 200,000
AI Inference Acceleration Algorithm Expert
AI Inference Acceleration Algorithm Expert

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 280,000
Staff Research Software Engineer: AI Inference & LLMs
Staff Research Software Engineer: AI Inference & LLMs

Google Inc. • Singapore

On-site
SGD 180,000 - 240,000
AI Infra & Live System Engineer (LLM Deployment)
AI Infra & Live System Engineer (LLM Deployment)

TIKTOK PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000