Intern Engineer – RL Post-Training for LLMs

Huawei Canada

Vancouver

On-site

CAD 58,000 - 104,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Huawei Canada is offering an immediate internship opportunity for an Intern Researcher, focusing on the development and optimization of RL post-training pipelines for Large Language Models. The role lets you collaborate with seasoned researchers on innovative LLM projects while enhancing your skills in reinforcement learning and deep learning.

Ideal candidates are pursuing a Master's or Ph.D. in Computer Science or AI, with proficiency in Python and PyTorch. The total target annual compensation ranges from $58,000 to $104,000 based on qualifications.

Qualifications

  • Strong background in machine learning, reinforcement learning, and deep learning.
  • Familiarity with Large Language Models and transformer architectures.
  • Hands-on experience with RL training algorithms is an asset.

Responsibilities

  • Develop and optimize RL post-training pipelines for LLMs.
  • Conduct experiments to improve model performance, reasoning, and alignment.
  • Collaborate with researchers on LLM projects.

Skills

Machine learning
Reinforcement learning
Deep learning
Python
PyTorch
LLM frameworks
Problem-solving
Communication skills

Education

Master or Ph.D. in Computer Science, AI, or related field

Tools

Hugging Face
DeepSpeed
vLLM
SGLang
RL frameworks (e.g., VeRL)

Job description

Huawei Canada has an immediate 6-12 months internship opening for an Intern Researcher.

About the team:

The Computing Data Application Acceleration Lab aims to create a leading global data analytics platform organized into three specialized teams using innovative programming technologies. This team focuses on full-stack innovations, including software-hardware co-design and optimizing data efficiency at both the storage and runtime layers. This team also develops next‑generation GPU architecture for gaming, cloud rendering, VR/AR, and Metaverse applications. One of the goals of this lab is to enhance algorithm performance and training efficiency across industries, fostering long‑term competitiveness.

About the job:
  • Develop and optimize RL post‑training pipelines for LLMs (e.g., GRPO, reward modeling).
  • Conduct experiments to improve model performance, reasoning, and alignment.
  • Build scalable training, evaluation, and data generation systems.
  • Collaborate with researchers and engineers on cutting‑edge LLM projects.
  • Stay current with advancements in RL, LLMs, and post‑training research.

The total target annual compensation (based on 2,080 hours per year) ranges from $58,000 to $104,000 depending on education, experience, and demonstrated expertise.

Job requirements
  • Enrolled as Master or Ph.D. student in Computer Science, AI, or related field.
  • Strong background in machine learning, reinforcement learning, and deep learning. Familiarity with Large Language Models, transformer architectures, and post‑training methods.
  • Proficiency in Python, PyTorch, and LLM frameworks.
  • Hands‑on experience with LLMs and RL training algorithms (e.g., GRPO) is an asset.
  • Familiarity with RL frameworks, such as VeRL.
  • Experience with open‑source LLM frameworks such as Hugging Face, DeepSpeed, vLLM, or SGLang is an asset.
  • Knowledge of domain‑specific languages used with AI accelerators.
  • Experience with distributed training frameworks, large‑scale experimentation, or LLM infrastructure is an asset.
  • Strong problem‑solving and communication skills.
Additional Information

Huawei Canada is committed to a fair, inclusive, and accessible recruitment process. If you require accommodation during any stage of the hiring process, please let us know and we will work with you to meet your needs.

All applications for this position are reviewed directly by our hiring team, we do not use artificial intelligence tools to screen or select candidates.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Intern Researcher - LLMs Agentic AI and RL
Intern Researcher - LLMs Agentic AI and RL

Huawei Canada • Markham

On-site
CAD 79,000 - 143,000
Intern Researcher - LLMs Agentic AI and RL
Intern Researcher - LLMs Agentic AI and RL

Huawei Technologies Canada Co., Ltd. • Markham

On-site
CAD 58,000 - 104,000
Machine Learning Researcher -LLM Agents & Efficient Deep Learning
Machine Learning Researcher -LLM Agents & Efficient Deep Learning

Huawei Canada • Montreal (administrative region)

On-site
CAD 106,000 - 156,000
Research Engineer - LLM Training & Alignment Systems
Research Engineer - LLM Training & Alignment Systems

Huawei Technologies Canada Co., Ltd. • Kingston

On-site
CAD 127,000 - 225,000
Intern Researcher – AI Foundation Model Training
Intern Researcher – AI Foundation Model Training

Huawei Canada • Markham

On-site
CAD 58,000 - 104,000
Engineer - ML & RL
Engineer - ML & RL

Huawei Technologies Canada Co., Ltd. • Edmonton

On-site
CAD 80,000 - 110,000
Intern Researcher – AI Foundation Model Training
Intern Researcher – AI Foundation Model Training

Huawei Technologies Canada Co., Ltd. • Markham

On-site
CAD 58,000 - 104,000
Intern Researcher - AI Agent Evaluation
Intern Researcher - AI Agent Evaluation

Huawei Canada • Markham

On-site
CAD 58,000 - 104,000
Intern Researcher - AI Agent
Intern Researcher - AI Agent

Huawei Technologies Canada Co., Ltd. • Markham

On-site
CAD 93,000 - 104,000
Intern Research Engineer - AI Agent Systems
Intern Research Engineer - AI Agent Systems

Huawei Canada • Markham

On-site
CAD 58,000 - 104,000