AI Inference & Training Acceleration Architect

Beijing Foreign Enterprise Management Consultants Co.,Ltd.

Singapore

On-site

SGD 180,000 - 280,000

Full time

3 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Huawei seeks a driven AI Training/Inference Acceleration Algorithm Expert to lead R&D on acceleration algorithms for large-scale models and multimodal domains. You will optimize for next-generation AI compute architectures and drive integration within proprietary frameworks to maximize compute efficiency.

Responsibilities include steering end-to-end implementation, exploring trends like long-sequence modeling and sparsity, and shaping algorithms to maintain Huawei's competitive AI computing

Qualifications

  • Master’s or PhD degree in Computer Science, Artificial Intelligence, Electronic Engineering, or a related technical discipline.
  • Solid technical foundation in mainstream Large Language Model (LLM) architectures and Mixture of Experts (MoE) models, with proven experience in large-scale model training, fine-tuning, or inference deployment.
  • Strong hands-on experience in AI acceleration technologies, including zero-redundancy optimizer techniques, distributed parallel strategies (TP/PP/SP/VP/DP), communication compression, and memory optimization.
  • Technical expertise in high-performance attention kernels and KV-cache compression is essential, alongside proficiency in model weight/activation quantization and sparsity-aware acceleration.
  • High proficiency in Python and deep familiarity with leading deep learning frameworks and high-performance acceleration libraries; expertise in hardware-level kernel optimization or native operator tuning is highly preferred.
  • Deep understanding of hardware-software co-design, with knowledge of hardware characteristics (memory bandwidth, compute cycles, and interconnects) and a track record of optimizing performance for models with 100B+ parameters.
  • Proven ability to perform complex AI model tuning and optimization in distributed or heterogeneous computing environments, translating hardware-specific features into significant algorithmic performance gains.

Responsibilities

  • Lead the research and development of AI training and inference acceleration algorithms for Agentic AI and Multimodal domains, specifically optimized for next-generation AI compute architectures to maximize compute efficiency and utilization.
  • Drive the end-to-end implementation of acceleration algorithms within proprietary AI frameworks and acceleration libraries, ensuring seamless integration and continuous iterative optimization based on real-world performance metrics.
  • Oversee technical insight and foresight within the training/inference domain, identifying emerging trends such as long-sequence modeling and sparsity to pre-plan and develop cutting-edge algorithms that ensure the sustained competitive advantage of AI computing platforms.

Skills

LLM architectures
Mixture of Experts
AI training
inference deployment
Python
kernel optimization
hardware-software co-design
quantization
distributed computing
high-performance DL frameworks

Education

Master’s degree in CS/AI
PhD in related field

Tools

PyTorch
TensorFlow
CUDA kernels

Job description

Huawei seeks a driven AI Training/Inference Acceleration Algorithm Expert to lead R&D on acceleration algorithms for large-scale models and multimodal domains. You will optimize for next-generation AI compute architectures and drive integration within proprietary frameworks to maximize compute efficiency.

Responsibilities include steering end-to-end implementation, exploring trends like long-sequence modeling and sparsity, and shaping algorithms to maintain Huawei's competitive AI computing

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Training & Inference Acceleration Architect
AI Training & Inference Acceleration Architect

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
AI Inference Acceleration Algorithm Expert
AI Inference Acceleration Algorithm Expert

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 280,000
AI Training/Inference Acceleration Algorithm Engineer
AI Training/Inference Acceleration Algorithm Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
AI Hardware Architect for Ultra-Efficient Inference
AI Hardware Architect for Ultra-Efficient Inference

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
AI & Large Language Model Expert - Intelligent Mobility
AI & Large Language Model Expert - Intelligent Mobility

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 300,000
Advanced Engineer (High-Efficiency AI Computing)
Advanced Engineer (High-Efficiency AI Computing)

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 210,000
AI Compute Architect: High-Performance Systems Lead
AI Compute Architect: High-Performance Systems Lead

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 260,000
AI Inference & Video Compression Engineer
AI Inference & Video Compression Engineer

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 120,000 - 180,000
Health insurance
AI Computing Architecture Researcher
AI Computing Architecture Researcher

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 150,000 - 260,000
AI and Large Language Model Expert
AI and Large Language Model Expert

Beijing Foreign Enterprise Management Consultants Co.,Ltd. • Singapore

On-site
SGD 180,000 - 300,000