AI Inference & Video Compression Architect

Beijing Foreign Enterprise Management Consultants Co.,Ltd.

Singapore

On-site

SGD 180,000 - 300,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Huawei in Singapore seeks an AI Inference & Compression Engineer to push the boundaries of LLM serving and video coding.

You will research compression algorithms, optimize KV caches and model quantization, and tackle memory bandwidth bottlenecks during autoregressive decoding.

In addition, you’ll develop AI-based video coding components, bridge R&D with production, and assess performance with PSNR, VMAF, latency, and throughput metrics.

Qualifications

  • Master’s or PhD in Computer Science, Electronic Engineering, Mathematics, or related fields (PhD preferred).
  • Solid understanding of video coding fundamentals including prediction, transform coding, quantization, and entropy coding with hands-on experience in standards such as H.265/HEVC, AV1, or H.266/VVC.
  • Strong understanding of Transformer architectures and attention mechanisms, as well as key performance bottlenecks in generative AI inference, particularly memory bandwidth constraints (“memory wall”).
  • Strong proficiency in Python and C/C++. Hands-on experience building, training, and modifying models using PyTorch, TensorFlow, etc.

Responsibilities

  • LLM Inference Acceleration. Research and develop advanced compression algorithms to accelerate LLM serving. Focus on KV cache optimization, model quantization, and resolving memory bandwidth bottlenecks during autoregressive decoding.
  • Classical Codec Development. Design and implement advanced video compression algorithms, focusing on improving RD performance, optimizing entropy coding, and enhancing quantization design for real-world applications.
  • AI-Based Media Coding. Develop and optimize AI-based video coding components, including AI-based loop filters, optical flow, and intelligent rate control.
  • Model Deployment & Fusion. Bridge the gap between AI research and production. Optimize deep learning models for efficient inference and ensure seamless integration of compression algorithms into deployment frameworks (e.g., vLLM).
  • Performance & Quality Evaluation. Conduct rigorous objective and subjective visual quality assessments such as PSNR and VMAF for video systems, as well as perplexity, zero-shot benchmarks, latency, and throughput analysis for LLM systems.

Skills

Python
C/C++
PyTorch
TensorFlow
Transformer architectures
Memory bandwidth

Education

Master’s or PhD in Computer Science/EE/Math

Tools

CUDA
HEVC/AV1/H.266 codecs

Job description

Huawei in Singapore seeks an AI Inference & Compression Engineer to push the boundaries of LLM serving and video coding.

You will research compression algorithms, optimize KV caches and model quantization, and tackle memory bandwidth bottlenecks during autoregressive decoding.

In addition, you’ll develop AI-based video coding components, bridge R&D with production, and assess performance with PSNR, VMAF, latency, and throughput metrics.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Inference & Compression Engineer
AI Inference & Compression Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 300,000
AI Compute Acceleration Engineer — Training & Inference
AI Compute Acceleration Engineer — Training & Inference

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 280,000
AI Training & Inference Acceleration Architect
AI Training & Inference Acceleration Architect

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
Video Coding R&D Engineer - Neural Compression & Standards
Video Coding R&D Engineer - Neural Compression & Standards

PANASONIC R&D CENTER SINGAPORE • Singapore

On-site
SGD 60,000 - 80,000
AI Hardware Architect - Low-Precision & Sparse Compute
AI Hardware Architect - Low-Precision & Sparse Compute

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 260,000
AI Inference Acceleration Algorithm Engineer
AI Inference Acceleration Algorithm Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 280,000
AI Training/Inference Acceleration Algorithm Engineer
AI Training/Inference Acceleration Algorithm Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
AI Engineer Intern — Real-Time Video Analytics (CCTV)
AI Engineer Intern — Real-Time Video Analytics (CCTV)

Mikomiko Pte. Ltd.mikomiko.ai • Singapore

On-site
SGD 90,000 - 150,000
High-Performance Backend Inference Engineer
High-Performance Backend Inference Engineer

Bytedance • Singapore

On-site
SGD 180,000 - 280,000
Vision AI Solutions Architect for VLM & Video AI
Vision AI Solutions Architect for VLM & Video AI

NVIDIA • Singapore

On-site
SGD 180,000 - 240,000