Distinguished Engineer - AI Computing System

Huawei Canada

Markham

On-site

CAD 172,000 - 230,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Huawei Canada is seeking a Distinguished Engineer in AI computing systems to lead development of training-cluster software frameworks and acceleration features for large model training and inference. You will guide low-precision training, parallel strategy tuning, and resource optimization to boost performance across AI clusters, collaborating with researchers and industry partners.

You will oversee a team, shape framework features for pre-training and integrated training/inference scenarios,

Qualifications

  • More than 5 years of R&D experience in large model training and optimization.
  • Proficient in large-model structures and optimization in areas such as LLMs, MoE, and multimodal learning.
  • Familiar with AI accelerators hardware and software-cores collaboration.
  • Experience in cluster computing and cloud computing for software architecture design.

Responsibilities

  • Lead and shape AI training frameworks for large model pre-training, post-training, and integrated training/inference.
  • Direct efforts in low-precision training, parallel strategy tuning, and training resource optimization.
  • Develop large model AI training frameworks, operator libraries, and acceleration features.
  • Identify academic resources, collaborate with researchers, and build long-term competitiveness in AI training clusters.
  • Cultivate a team of technical experts and key technical backbone.

Skills

Large model training
AI framework design
Software optimization
GPU/AI accelerators
Cluster computing & scheduling
Team leadership
Research ability

Education

Advanced degree in AI/CS/related field

Tools

PyTorch/TensorFlow
CUDA/CUDA-X
Deep learning libraries

Job description

Huawei Canada has an immediate permanent opening for a Distinguished Engineer - AI Computing System

About the team

The Advanced Computing and Storage Lab, currently a part of the Vancouver Research Centre, aims to explore adaptive computing system architectures to address the challenges posed by flexible and variable application loads in the future. It assists in ensuring the stability and quality of training clusters, constructs dynamic cluster configuration strategy solvers, and establishes precision control systems to create stable and efficient computing power clusters. One of the lab's goals is to focus on key industry AI application scenarios such as large model training/inference, based on key technologies like low-precision training, multi-modal training, and reinforcement learning, responsible for bottleneck analysis and the design and development of optimization solutions, thereby improving training and inference performance as well as usability.


About the job


  • As a leading expert in the industry in the field of training cluster software frameworks and technologies, gain insights into the evolution direction of industry AI large model training frameworks and key features. Plan and layout AI frameworks and software features for scenarios such as large model pre-training, post-training, and integrated training and inference, building key capabilities for the company's training cluster software framework.

  • Focusing on the company's large model training optimization field, lead the team to build key technologies such as low-precision training, parallel strategy tuning, and training resource optimization, promoting the commercial implementation of large model perception optimization-related technologies.

  • Focusing on the company's training servers and super nodes and other products, lead the team to build large model AI training frameworks, operator libraries, acceleration libraries, and other software frameworks and acceleration features, fully leveraging system engineering and software-hardware collaboration capabilities to enhance AI cluster computing efficiency.

  • Identify high-quality academic resources in the direction of large model training, collaborate with domain experts and scholars on projects, layout related standards and patents, support the company's continuous innovation in the training cluster field, and build long-term competitiveness in the AI training cluster direction.

  • Cultivate a team of technical experts and key technical backbone in the direction of AI training cluster frameworks and software optimization.


The base salary for this position ranges from $172,000 to $230,000 depending on education, experience and demonstrated expertise.


Job requirements

About The Ideal Candidate


  • Major in artificial intelligence, computer science, software, automation, physics, mathematics, electronics, microelectronics, information technology, or related fields, with more than 5 years of R&D experience in large model training and optimization.

  • Proficient in common model structures of large models such as Deepseek and Llama, with deep technical expertise in large model training and inference optimization in fields like LLM, MoE, and multimodal learning.

  • Familiar with the hardware architecture and programming systems of AI accelerators such as GPU and NPU, with experience in optimizing AI systems with software-hardware-cores collaboration.

  • Familiar with cluster computing and cloud computing fields, with experience in software architecture design for cluster scheduling.

  • Enjoys research, has strong learning ability, good communication skills, and teamwork ability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Researcher - AI Computing System
Researcher - AI Computing System

Huawei Canada • Vancouver

On-site
CAD 106,000 - 205,000
Senior Technical VP - AI & Efficient Deep Learning
Senior Technical VP - AI & Efficient Deep Learning

Huawei Canada • Markham

On-site
CAD 150,000 - 200,000
Co-op Researcher - AI Computing System
Co-op Researcher - AI Computing System

Huawei Canada • Vancouver

On-site
CAD 76,000 - 109,000
Intern Researcher - AI Computing System
Intern Researcher - AI Computing System

Huawei Canada • Vancouver

On-site
CAD 78,000 - 150,000
Senior Researcher – Hardware Efficient AI Foundation Model Training
Senior Researcher – Hardware Efficient AI Foundation Model Training

Huawei Canada • Markham

On-site
CAD 127,000 - 225,000
Senior Engineer-Cloud AI Infrastructure
Senior Engineer-Cloud AI Infrastructure

Huawei Canada • Markham

On-site
CAD 172,000 - 306,000
Research Engineer - AI Workload & Systems
Research Engineer - AI Workload & Systems

Huawei Technologies Canada Co., Ltd. • Markham

On-site
CAD 178,000 - 316,000
Intern Researcher – AI Foundation Model Training
Intern Researcher – AI Foundation Model Training

Huawei Technologies Canada Co., Ltd. • Markham

On-site
CAD 58,000 - 104,000
Technical VP - AI Data Platform
Technical VP - AI Data Platform

Huawei Canada • Markham

On-site
CAD 180,000 - 240,000
Senior AI Technology Strategy Expert
Senior AI Technology Strategy Expert

Huawei Canada • Markham

On-site
CAD 110,000 - 190,000