AI Performance Engineer – GPU & ROCm

Sapphire Stream Technology

Toronto

On-site

CAD 90,000 - 140,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Sapphire Stream Technology in Toronto, Ontario, is seeking a GPU AI Performance Engineer to optimize and deploy AI workloads using ROCm. You will focus on increasing training speed, reducing memory usage, and improving multi-GPU scalability across Linux and cloud environments.

You will work with GPU software, AI frameworks like PyTorch, TensorFlow and ONNX Runtime, profile with ROCm tools, and collaborate with hardware, compiler, and infrastructure teams to deliver robust performance

Qualifications

  • Bachelor’s or master’s degree in CS, CE, EE, or related field.
  • Proficient in C++ and Python.
  • Experience with Linux and GPU computing.
  • Hands-on ROCm, HIP, CUDA, or similar GPU technologies.
  • Understanding of GPU architecture and memory management.
  • Experience with PyTorch, TensorFlow, ONNX Runtime, or similar AI frameworks.
  • Experience profiling and optimizing GPU applications.
  • Knowledge of AI model training, inference, and performance optimization.

Responsibilities

  • Improve training speed, inference latency, memory usage, and multi-GPU scalability.
  • Work with large language models, vision models, multimodal AI, and generative AI.
  • Develop and troubleshoot applications using ROCm and HIP.
  • Work with PyTorch, TensorFlow, ONNX Runtime, or MIGraphX (and related frameworks).
  • Profile GPU applications and identify performance bottlenecks.
  • Apply memory optimization, mixed precision, and kernel tuning techniques.
  • Support AI deployment on Linux, cloud, edge, and containerized platforms.
  • Develop performance benchmarks and automated validation tests.
  • Collaborate with hardware, compiler, runtime, framework, and infrastructure teams.

Skills

C++
Python
Linux
ROCm
CUDA
GPU architecture
PyTorch
TensorFlow
ONNX Runtime
Profiling
Performance optimization

Education

Bachelor's or Master's degree in CS/CE/EE

Tools

ROCm
HIP
CUDA
PyTorch
TensorFlow
ONNX Runtime

Job description

We are looking for a GPU AI Performance Engineer to optimize and deploy AI and machine-learning workloads using ROCm.

You will improve the speed, memory efficiency, and scalability of AI training and inference workloads. You will also work with GPU software, AI frameworks, profiling tools, and cloud or containerized environments.

Key Responsibilities
  • Improve training speed, inference latency, memory usage, and multi-GPU scalability.
  • Work with large language models, vision models, multimodal AI, and generative AI.
  • Develop and troubleshoot applications using ROCm and HIP.
  • Work with PyTorch, TensorFlow, ONNX Runtime, vLLM, SGLang, or MIGraphX.
  • Profile GPU applications and identify performance bottlenecks.
  • Apply mixed precision, quantization, kernel tuning, operator fusion, and memory optimization.
  • Support AI deployment on Linux, cloud, edge, and containerized platforms.
  • Develop performance benchmarks and automated validation tests.
  • Collaborate with hardware, compiler, runtime, framework, and infrastructure teams.
Required Qualifications
  • Bachelor’s or master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
  • Strong programming skills in C++ and Python.
  • Experience with Linux and GPU computing.
  • Hands‑on experience with ROCm, HIP, CUDA, or similar GPU technologies.
  • Understanding of GPU architecture, parallel programming, and memory management.
  • Experience with PyTorch, TensorFlow, ONNX Runtime, or similar AI frameworks.
  • Experience profiling and optimizing GPU applications.
  • Knowledge of AI model training, inference, and performance optimization.
Preferred Qualifications
  • Experience optimizing large language models or generative AI applications.
  • Experience with vLLM, SGLang, or distributed inference.
  • Familiarity with LLVM, MLIR, or compiler technologies.
  • Experience migrating workloads between CUDA and HIP.
  • Knowledge of Kubernetes, containers, and distributed AI systems.
  • Contributions to ROCm or other open-source GPU projects.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

HPC Specialist
HPC Specialist

DRW Holdings, LLC. • Montreal (administrative region)

On-site
CAD 100,000 - 130,000
AI Model, Framework, and GPU Software Engineer – Agentic AI
AI Model, Framework, and GPU Software Engineer – Agentic AI

AMD • Markham

On-site
CAD 110,000 - 160,000
Senior Technical VP - AI & Efficient Deep Learning
Senior Technical VP - AI & Efficient Deep Learning

Huawei Canada • Markham

On-site
CAD 150,000 - 200,000
Senior Researcher – GPU, AI & Hardware Architecture
Senior Researcher – GPU, AI & Hardware Architecture

Huawei Technologies Canada Co., Ltd. • Edmonton

On-site
CAD 120,000 - 170,000
GPU Cloud Platform Engineer
GPU Cloud Platform Engineer

Yotta Labs • Canada

On-site
CAD 90,000 - 120,000
Flexible work environment
Cutting-edge technology projects
Collaborative team culture
Senior Researcher – GPU, AI & Hardware Architecture
Senior Researcher – GPU, AI & Hardware Architecture

Huawei Canada • Edmonton

On-site
CAD 90,000 - 140,000
AI Systems Engineer – AI Model (Training & Inference)
AI Systems Engineer – AI Model (Training & Inference)

AMD • Markham

On-site
CAD 140,000 - 190,000
Compute Solution Architect
Compute Solution Architect

Mistral • Montreal (administrative region)

On-site
CAD 110,000 - 150,000
Site Reliability Engineer, AI/ML Infrastructure
Site Reliability Engineer, AI/ML Infrastructure

Boson AI • Toronto

On-site
CAD 100,000 - 130,000
Senior Researcher - GPU Applications Architecture
Senior Researcher - GPU Applications Architecture

Huawei Canada • Markham

On-site
CAD 90,000 - 120,000