AI Inference Acceleration Algorithm Expert

Beijing Foreign Enterprise Management Consultants Co.,Ltd.

Singapore

On-site

SGD 180,000 - 280,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Huawei seeks a driven AI Training/Inference Acceleration Algorithm Expert to lead R&D on acceleration algorithms for large-scale models and multimodal domains. You will optimize for next-generation AI compute architectures and drive integration within proprietary frameworks to maximize compute efficiency.

Responsibilities include steering end-to-end implementation, exploring trends like long-sequence modeling and sparsity, and shaping algorithms to maintain Huawei's competitive AI computing

Qualifications

  • Master’s or PhD degree in Computer Science, Artificial Intelligence, Electronic Engineering, or a related technical discipline.
  • Solid technical foundation in mainstream Large Language Model (LLM) architectures and Mixture of Experts (MoE) models, with proven experience in large-scale model training, fine-tuning, or inference deployment.
  • Strong hands-on experience in AI acceleration technologies, including zero-redundancy optimizer techniques, distributed parallel strategies (TP/PP/SP/VP/DP), communication compression, and memory optimization.
  • Technical expertise in high-performance attention kernels and KV-cache compression is essential, alongside proficiency in model weight/activation quantization and sparsity-aware acceleration.
  • High proficiency in Python and deep familiarity with leading deep learning frameworks and high-performance acceleration libraries; expertise in hardware-level kernel optimization or native operator tuning is highly preferred.
  • Deep understanding of hardware-software co-design, with knowledge of hardware characteristics (memory bandwidth, compute cycles, and interconnects) and a track record of optimizing performance for models with 100B+ parameters.
  • Proven ability to perform complex AI model tuning and optimization in distributed or heterogeneous computing environments, translating hardware-specific features into significant algorithmic performance gains.

Responsibilities

  • Lead the research and development of AI training and inference acceleration algorithms for Agentic AI and Multimodal domains, specifically optimized for next-generation AI compute architectures to maximize compute efficiency and utilization.
  • Drive the end-to-end implementation of acceleration algorithms within proprietary AI frameworks and acceleration libraries, ensuring seamless integration and continuous iterative optimization based on real-world performance metrics.
  • Oversee technical insight and foresight within the training/inference domain, identifying emerging trends such as long-sequence modeling and sparsity to pre-plan and develop cutting-edge algorithms that ensure the sustained competitive advantage of AI computing platforms.

Skills

LLM architectures
Mixture of Experts
AI training
inference deployment
Python
kernel optimization
hardware-software co-design
quantization
distributed computing
high-performance DL frameworks

Education

Master’s degree in CS/AI
PhD in related field

Tools

PyTorch
TensorFlow
CUDA kernels

Job description

On behalf of Huawei, a world-renowned information and communication technology company, we are seeking passionate and talented individuals to join our team as AI Training/Inference Acceleration Algorithm Expert.

Job Description:

  • Lead the research and development of AI training and inference acceleration algorithms for Agentic AI and Multimodal domains, specifically optimized for next-generation AI compute architectures to maximize compute efficiency and utilization.
  • Drive the end-to-end implementation of acceleration algorithms within proprietary AI frameworks and acceleration libraries, ensuring seamless integration and continuous iterative optimization based on real-world performance metrics.
  • Oversee technical insight and foresight within the training/inference domain, identifying emerging trends such as long-sequence modeling and sparsity to pre-plan and develop cutting-edge algorithms that ensure the sustained competitive advantage of AI computing platforms.

Skills / Qualifications:

  • Master’s or PhD degree in Computer Science, Artificial Intelligence, Electronic Engineering, or a related technical discipline.
  • Solid technical foundation in mainstream Large Language Model (LLM) architectures and Mixture of Experts (MoE) models, with proven experience in large-scale model training, fine-tuning, or inference deployment.
  • Strong hands-on experience in AI acceleration technologies, including zero-redundancy optimizer techniques, distributed parallel strategies (TP/PP/SP/VP/DP), communication compression, and memory optimization.
  • Technical expertise in high-performance attention kernels and KV-cache compression is essential, alongside proficiency in model weight/activation quantization and sparsity-aware acceleration.
  • High proficiency in Python and deep familiarity with leading deep learning frameworks and high-performance acceleration libraries; expertise in hardware-level kernel optimization or native operator tuning is highly preferred.
  • Deep understanding of hardware-software co-design, with knowledge of hardware characteristics (memory bandwidth, compute cycles, and interconnects) and a track record of optimizing performance for models with 100B+ parameters.
  • Proven ability to perform complex AI model tuning and optimization in distributed or heterogeneous computing environments, translating hardware-specific features into significant algorithmic performance gains.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Training/Inference Acceleration Algorithm Engineer
AI Training/Inference Acceleration Algorithm Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
AI Inference & Training Acceleration Architect
AI Inference & Training Acceleration Architect

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 280,000
Advanced Engineer (High-Efficiency AI Computing)
Advanced Engineer (High-Efficiency AI Computing)

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
AI Computing Architecture Researcher
AI Computing Architecture Researcher

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 260,000
AI Training & Inference Acceleration Architect
AI Training & Inference Acceleration Architect

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
Advanced Engineer (High-Efficiency AI Computing)
Advanced Engineer (High-Efficiency AI Computing)

PERSOL SINGAPORE PTE. LTD. • Singapore

On-site
SGD 180,000 - 300,000
AI and Large Language Model Expert
AI and Large Language Model Expert

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 300,000
AI Inference & Compression Engineer
AI Inference & Compression Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 120,000 - 180,000
Health insurance
AI Hardware Architect for Ultra-Efficient Inference
AI Hardware Architect for Ultra-Efficient Inference

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
AI Computing Architecture Researcher
AI Computing Architecture Researcher

PERSOL SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000