AI Inference Acceleration Algorithm Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd.

Singapore

On-site

SGD 180,000 - 280,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Huawei in Singapore seeks an AI Training/Inference Acceleration Algorithm Engineer to lead R&D on acceleration for large-scale models and multimodal domains. You will optimize for next-generation compute architectures, implement algorithms in proprietary and open frameworks, and collaborate across hardware and software teams to achieve peak performance.

Strong track record in LLMs, MoE, and distributed training is required, with expertise in kernel optimization, quantization, memory management,

Qualifications

  • Master's or PhD degree in Computer Science, Artificial Intelligence, Electronic Engineering, or a related discipline.
  • Solid foundation in mainstream LLM architectures and MoE models with experience in large-scale model training, fine-tuning, or inference deployment.
  • Strong hands-on experience in AI acceleration technologies including zero-redundancy optimizers, distributed parallel strategies, communication compression, and memory optimization.
  • Technical expertise in high-performance attention kernels and KV-cache compression; experience with weight/activation quantization and sparsity-aware acceleration.
  • Proficiency in Python and familiarity with deep learning frameworks; hardware-level kernel optimization or native operator tuning is preferred.
  • Understanding of hardware-software co-design and hardware characteristics; track record optimizing performance for 100B+ parameter models.
  • Ability to perform complex AI model tuning and optimization in distributed or heterogeneous environments with hardware-aware algorithmic gains.

Responsibilities

  • Lead the research and development of AI training and inference acceleration algorithms for Agentic AI and Multimodal domains, optimized for next-generation AI compute architectures.
  • Drive end-to-end implementation of acceleration algorithms within proprietary AI frameworks and libraries, ensuring seamless integration and iterative optimization based on performance metrics.
  • Oversee technical insight and foresight in training/inference, identifying trends like long-sequence modeling and sparsity to maintain competitive AI computing platforms.

Skills

LLM architectures
MoE models
AI acceleration
Python
Deep learning frameworks
Kernel optimization
Hardware-software co-design
Distributed training
Quantization
Sparsity-aware acceleration

Education

Master's or PhD in CS/AI/EE

Tools

TensorFlow/PyTorch

Job description

On behalf of Huawei, a world-renowned information and communication technology company, we are seeking passionate and talented individuals to join our team as AI Training/Inference Acceleration Algorithm Engineer.

Job Description:

  • Lead the research and development of AI training and inference acceleration algorithms for Agentic AI and Multimodal domains, specifically optimized for next-generation AI compute architectures to maximize compute efficiency and utilization.
  • Drive the end-to-end implementation of acceleration algorithms within proprietary AI frameworks and acceleration libraries, ensuring seamless integration and continuous iterative optimization based on real-world performance metrics.
  • Oversee technical insight and foresight within the training/inference domain, identifying emerging trends such as long-sequence modeling and sparsity to pre-plan and develop cutting-edge algorithms that ensure the sustained competitive advantage of AI computing platforms.

Skills / Qualifications:

  • Master’s or PhD degree in Computer Science, Artificial Intelligence, Electronic Engineering, or a related technical discipline.
  • Solid technical foundation in mainstream Large Language Model (LLM) architectures and Mixture of Experts (MoE) models, with proven experience in large-scale model training, fine-tuning, or inference deployment.
  • Strong hands-on experience in AI acceleration technologies, including zero-redundancy optimizer techniques, distributed parallel strategies (TP/PP/SP/VP/DP), communication compression, and memory optimization.
  • Technical expertise in high-performance attention kernels and KV-cache compression is essential, alongside proficiency in model weight/activation quantization and sparsity-aware acceleration.
  • High proficiency in Python and deep familiarity with leading deep learning frameworks and high-performance acceleration libraries; expertise in hardware-level kernel optimization or native operator tuning is highly preferred.
  • Deep understanding of hardware-software co-design, with knowledge of hardware characteristics (memory bandwidth, compute cycles, and interconnects) and a track record of optimizing performance for models with 100B+ parameters.
  • Proven ability to perform complex AI model tuning and optimization in distributed or heterogeneous computing environments, translating hardware-specific features into significant algorithmic performance gains.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Training/Inference Acceleration Algorithm Engineer
AI Training/Inference Acceleration Algorithm Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
AI Compute Acceleration Engineer — Training & Inference
AI Compute Acceleration Engineer — Training & Inference

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 280,000
AI Training & Inference Acceleration Architect
AI Training & Inference Acceleration Architect

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
Advanced Engineer (High-Efficiency AI Computing)
Advanced Engineer (High-Efficiency AI Computing)

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 260,000
AI Computing Architecture Researcher
AI Computing Architecture Researcher

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 120,000 - 180,000
AI and Large Language Model Expert
AI and Large Language Model Expert

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 320,000
Senior Algorithm Engineer (Large Models & Autonomous Driving)
Senior Algorithm Engineer (Large Models & Autonomous Driving)

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 190,000
AI Engineer Intern
AI Engineer Intern

Tencent • Singapore

On-site
SGD 80,000 - 100,000
AI Inference & Compression Engineer
AI Inference & Compression Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 300,000
AI Hardware Architect - Low-Precision & Sparse Compute
AI Hardware Architect - Low-Precision & Sparse Compute

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 260,000