AI Inference & Compression Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd.

Singapore

On-site

SGD 120,000 - 180,000

Full time

29 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Health insurance

Job summary

Huawei is seeking an AI Inference & Compression Engineer in Singapore to advance LLM inference acceleration, develop advanced video codecs, and optimize AI-based media coding. The role includes deploying optimized models and integrating compression techniques into deployment frameworks such as vLLM.

The candidate will work on performance and quality evaluation, including PSNR, VMAF, perplexity, latency, and throughput, ensuring robust, scalable AI/video systems.

Qualifications

  • Master’s or PhD in Computer Science, Electronics, Mathematics, or related fields (PhD preferred).
  • Solid understanding of video coding fundamentals incl. prediction, transform coding, quantization, entropy coding (H.265/HEVC, AV1, H.266/VVC).
  • Strong understanding of Transformer architectures and attention mechanisms, memory bandwidth constraints in AI inference.
  • Proficiency in Python and C/C++, with hands-on experience in PyTorch, TensorFlow, etc.

Responsibilities

  • LLM Inference Acceleration: research and develop compression algorithms to accelerate LLM serving, KV cache optimization, decoding bottlenecks.
  • Classical Codec Development: improve RD performance, optimize entropy coding, enhance quantization for real-world apps.
  • AI-Based Media Coding: develop AI-based loop filters, optical flow, and rate control.
  • Model Deployment & Fusion: bridge AI research and production, optimize models for inference and integrate compression into frameworks like vLLM.
  • Performance & Quality Evaluation: assess PSNR, VMAF, perplexity, latency, and throughput for AI/video systems.

Skills

Python
C/C++
PyTorch
TensorFlow
Transformer
Memory bandwidth
H265/HEVC
AV1
H266/VVC

Education

Master's or PhD in CS/EE/Math

Tools

vLLM

Job description

On behalf of Huawei, a world-renowned information and communication technology company, we are seeking passionate and talented individuals to join our team as AI Inference & Compression Engineer.

Key Responsibilities
  • LLM Inference Acceleration. Research and develop advanced compression algorithms to accelerate LLM serving. Focus on KV cache optimization, model quantization, and resolving memory bandwidth bottlenecks during autoregressive decoding.
  • Classical Codec Development. Design and implement advanced video compression algorithms, focusing on improving Rate–Distortion (RD) performance, optimizing entropy coding, and enhancing quantization design for real-world applications.
  • AI-Based Media Coding. Develop and optimize AI-based video coding components, including AI-based loop filters, optical flow, and intelligent rate control.
  • Model Deployment & Fusion. Bridge the gap between AI research and production. Optimize deep learning models for efficient inference and ensure seamless integration of compression algorithms into deployment frameworks (e.g., vLLM).
  • Performance & Quality Evaluation. Conduct rigorous objective and subjective visual quality assessments such as PSNR and VMAF for video systems, as well as perplexity, zero-shot benchmarks, latency, and throughput analysis for LLM systems.
Required Qualifications
  • Master’s or PhD in Computer Science, Electronic Engineering, Mathematics, or related fields (PhD preferred).
  • Solid understanding of video coding fundamentals including prediction, transform coding, quantization, and entropy coding with hands‑on experience in standards such as H.265/HEVC, AV1, or H.266/VVC.
  • Strong understanding of Transformer architectures and attention mechanisms, as well as key performance bottlenecks in generative AI inference, particularly memory bandwidth constraints (“memory wall”).
  • Strong proficiency in Python and C/C++. Hands‑on experience building, training, and modifying models using PyTorch, TensorFlow, etc.
Preferred Qualifications
  • ISP Knowledge. Familiarity with Image Signal Processing flow, such as demosaicing, denoising, and tone mapping.
  • Image Processing. Experience in computer vision-based image enhancement (e.g., de‑blurring, artifact removal, or HDR).
  • Hardware Optimization. Knowledge of SIMD, CUDA, or other hardware acceleration techniques for video and tensor processing.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Inference & Video Compression Engineer
AI Inference & Video Compression Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 120,000 - 180,000
Health insurance
AI Inference Acceleration Algorithm Expert
AI Inference Acceleration Algorithm Expert

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 280,000
AI Training/Inference Acceleration Algorithm Engineer
AI Training/Inference Acceleration Algorithm Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
Intelligent Video Processing Algorithm Engineer
Intelligent Video Processing Algorithm Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 120,000 - 240,000
Image Signal Processing Algorithm Engineer
Image Signal Processing Algorithm Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 100,000 - 150,000
Advanced Engineer (High-Efficiency AI Computing)
Advanced Engineer (High-Efficiency AI Computing)

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
Intelligent Video Processing Algorithm Engineer
Intelligent Video Processing Algorithm Engineer

PERSOL SINGAPORE PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
AI Computing Architecture Researcher
AI Computing Architecture Researcher

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 260,000
Embedded LLM Systems Engineer
Embedded LLM Systems Engineer

Desay SV • Singapore

On-site
SGD 120,000 - 180,000
Senior Algorithm Engineer (Multimodal Generation)
Senior Algorithm Engineer (Multimodal Generation)

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 110,000 - 170,000