HPC AI Engineer, Frontier, NSCC

A*STAR Research

Canada

On-site

CAD 80,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading research institution in Canada is seeking an HPC AI Engineer to support large-scale AI applications using a new supercomputer system. You will collaborate with researchers, optimise AI workloads, and develop best practices for HPC software. A Bachelor's degree in computer science or engineering is required along with programming skills in Python and knowledge of AI frameworks. The role involves engagement with diverse communities and includes responsibilities in advising users and developing utilities.

Qualifications

  • Proven working knowledge of models and algorithms in generative models, computer vision, or AI for Science.
  • Ideally, 3 years of experience in developing codes for AI training and inference.
  • Familiar with GPU architectures and programming.

Responsibilities

  • Provide HPC and scientific domain advice to users of NSCC systems.
  • Support and optimise large-scale AI application workloads.
  • Design, develop and implement HPC software best practices.

Skills

Python Programming
Algorithms in AI
AI Performance Optimisation
AI Frameworks (e.g., PyTorch, Tensorflow)
C/C++ Programming
Linux Environment

Education

Bachelor degree in computer science, computer engineering, or relevant areas

Tools

HPC job schedulers
Container technologies
S3 Object Storage

Job description

Requisition ID: 593

Posting Start Date: 01/04/2026

About the role

As our HPC AI Engineer, you will be a key expert supporting researchers in leveraging our new supercomputer system for large-scale artificial intelligence. You will support and optimise massive AI application workloads, working with performance engineers to profile AI applications and establish best practices. Your work will directly enable national-scale projects in multimodal AI, healthcare, and AI for Science.

RESPONSIBILITIES
  • Provide HPC and scientific domain advice to users of NSCC systems.
  • Engage and collaborate with new researchers, communities, and disciplines with computationally intensive requirements.
  • Support and optimise large-scale AI application workloads.
  • Work with HPC performance engineers to profile and build performance models of the AI applications and workflows.
  • Design, develop and implement HPC software best practices for AI applications and workflows.
  • Assist in the planning and design of future HPC systems, including benchmarking AI workloads on various platforms and recommending the most suitable architecture for the research community.
  • Analyse system and user job data for efficient resource allocation and management.
  • Develop HPC utilities, dashboards and automated testing tools for NSCC HPC systems.
  • Develop HPC user and best practice guides for NSCC HPC systems.
  • Get up-to-date with scientific domain research development, HPC system and software technology
QUALIFICATIONS
  • Bachelor degree in the field of computer science, computer engineering, or other relevant areas.
  • Proven working knowledge of models and algorithms in at least one area of generative models, computer vision, graph neural networks, or AI for Science applications.
  • Ideally, 3 years of experience in developing codes for AI training and inference.
  • Experience in setting up AI software stacks, familiar with diversified AI software stacks.
  • Good knowledge in AI application performance optimisation and troubleshooting.
  • Strong programming skills in Python; familiar with C/C++ programming is a plus.
  • Familiar with the working and using of AI frameworks (e.g. PyTorch, Tensorflow, JAX) for research.
  • Familiar with GPU architectures and programming is highly desired.
  • Familiar with Linux environment, scripting languages, profiler and debugger tools.
  • Familiar with HPC job schedulers and container technologies.
  • Familiar with object storage (S3); familiar with HPC storage (Lustre) is a plus.
  • Demonstrated team player with strong problem-solving skills.
  • Demonstrated effective communication skills including the ability to articulate technical concepts to a diverse range of audiences.
  • Demonstrated ability and willingness to contribute novel ideas and approaches in support of the research community
  • Demonstrated passion for continuous learning and exploring new technologies or domains.

The above eligibility criteria are not exhaustive. A*STAR may include additional selection criteria based on its prevailing recruitment policies. These policies may be amended from time to time without notice. We regret that only shortlisted candidates will be notified.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC AI Engineer, Frontier, NSCC
HPC AI Engineer, Frontier, NSCC

A*STAR Research • Dartmouth

On-site
CAD 80,000 - 100,000
Researcher - AI Computing System
Researcher - AI Computing System

Huawei Canada • Vancouver

On-site
CAD 106,000 - 205,000
Intern Researcher - AI Computing System
Intern Researcher - AI Computing System

Huawei Canada • Vancouver

On-site
CAD 78,000 - 150,000
Co-op Researcher - AI Computing System
Co-op Researcher - AI Computing System

Huawei Canada • Vancouver

On-site
CAD 76,000 - 109,000
Senior Engineer-Cloud AI Infrastructure
Senior Engineer-Cloud AI Infrastructure

Huawei Canada • Markham

On-site
CAD 172,000 - 306,000
Senior Principal Researcher - AI for Science
Senior Principal Researcher - AI for Science

Huawei Canada • Markham

On-site
CAD 100,000 - 140,000
Distinguished Engineer - AI Computing System
Distinguished Engineer - AI Computing System

Huawei Canada • Markham

On-site
CAD 172,000 - 230,000
Research Engineer - AI Workload & Systems
Research Engineer - AI Workload & Systems

Huawei Technologies Canada Co., Ltd. • Markham

On-site
CAD 178,000 - 316,000
Site Reliability Engineer, AI/ML Infrastructure
Site Reliability Engineer, AI/ML Infrastructure

Boson AI • Toronto

On-site
CAD 100,000 - 130,000
Senior Researcher – Hardware Efficient AI Foundation Model Training
Senior Researcher – Hardware Efficient AI Foundation Model Training

Huawei Canada • Markham

On-site
CAD 127,000 - 225,000