AI Infrastructure Engineer — GPU Clusters & HPC

Veeda AI

Toronto

On-site

CAD 120,000 - 170,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Veeda AI in Toronto, Ontario seeks a Member of Technical Staff to advance AI infrastructure for multimodal foundation world models. You will own GPU cluster operations, security, data transmission, and end-to-end production observability.

The role emphasizes hands-on HPC expertise, kernel and networking tuning, IaC automation, and collaboration with model and data teams to scale compute and data resources.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in HPC or infrastructure engineering.
  • Strong troubleshooting skills below the framework layer: low-level Linux networking, kernel and PCIe/NUMA tuning, hardware diagnostics, and network or distributed storage (NFS, NVMe-oF, Lustre, Ceph).
  • Proficient in automation and infrastructure-as-code tools (e.g., Ansible, Terraform, Helm, Python/Bash scripting).
  • Experience supporting distributed deep learning workloads (PyTorch/NCCL, Ray, DeepSpeed) closely enough to tell an infrastructure fault from a model bug.

Responsibilities

  • Design, deploy, and operate Slurm-on-Kubernetes GPU clusters.
  • Own the compute, interconnect, storage, and scheduler resources for GPU clusters.
  • Design, deploy, and operate identity provisioning and access control for sources, clusters, data storages, and dev tools.
  • Design and maintain distributed data transmission and caching services for timely data delivery to clusters.
  • Build telemetry pipelines (Prometheus, Grafana, DCGM, BMC/Redfish) to catch errors and diagnose issues.
  • Lead PoC rounds with vendors for compute and data resources.
  • Lead projects through the full software lifecycle including CI/CD and on-call support.

Skills

HPC infrastructure
Linux troubleshooting
Automation
Python scripting
PyTorch/NCCL/Ray/DeepSpeed
Slurm
Kubernetes
GPU hardware

Education

Bachelor's degree in Computer Science or Computer Engineering

Tools

Ansible
Terraform
Helm
Python/Bash scripting
NVIDIA GPU hardware
Slurm/Kubernetes cluster management

Job description

Veeda AI in Toronto, Ontario seeks a Member of Technical Staff to advance AI infrastructure for multimodal foundation world models. You will own GPU cluster operations, security, data transmission, and end-to-end production observability.

The role emphasizes hands-on HPC expertise, kernel and networking tuning, IaC automation, and collaboration with model and data teams to scale compute and data resources.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote: Technical Lead, GPU Infra & HPC Platform
Remote: Technical Lead, GPU Infra & HPC Platform

Jobgether • Canada

Remote
CAD 150,000 - 230,000
100% remote position
Lead architecture and delivery of GPU/
International and distributed team
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Toronto

On-site
CAD 120,000 - 170,000
Site Reliability Engineer, AI/ML Infrastructure
Site Reliability Engineer, AI/ML Infrastructure

Boson AI • Toronto

On-site
CAD 100,000 - 130,000
Senior AI Inference Systems Engineer
Senior AI Inference Systems Engineer

NVIDIA Corporation • Toronto

Hybrid
CAD 170,000 - 275,000
Equity
Benefits
GPU Infrastructure Engineer for AI Inference & Serving
GPU Infrastructure Engineer for AI Inference & Serving

P2P • Montreal (administrative region)

On-site
CAD 90,000 - 120,000
Senior Technical VP - AI & Efficient Deep Learning
Senior Technical VP - AI & Efficient Deep Learning

Huawei Canada • Markham

On-site
CAD 150,000 - 200,000
Senior Production AI Systems Engineer
Senior Production AI Systems Engineer

Jaide Health • Toronto

Hybrid
CAD 140,000 - 210,000
Lunch stipend
Health and dental benefits
RRSP matching
+5
GPU Infra Architect for AI Inference
GPU Infra Architect for AI Inference

DRW Holdings, LLC. • Montreal (administrative region)

On-site
CAD 100,000 - 130,000
Senior AI/GPU Architect - Edge Compute & Hardware Advantage
Senior AI/GPU Architect - Edge Compute & Hardware Advantage

Huawei Technologies Canada Co., Ltd. • Edmonton

On-site
CAD 120,000 - 170,000
Principal AI Infrastructure Engineer
Principal AI Infrastructure Engineer

Advanced Micro Devices • Vancouver

On-site
CAD 140,000 - 190,000
Benefits at a glance