Senior AI Infrastructure Engineer — HPC & Compute Clusters

Veeda AI

Toronto

On-site

CAD 100,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Veeda AI in Toronto is building the next generation of multimodal foundation world models for Physical AI. This role focuses on GPU cluster orchestration and large-scale HPC infrastructure to support multi-node distributed training.

You will optimize Slurm and Kubernetes, manage high-speed networks like InfiniBand and NVLink, and build observability and reliability pipelines to keep thousands of GPUs saturated.

Qualifications

  • You have a Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in high-performance computing (HPC) or infrastructure engineering.
  • You have deep hands-on expertise administering Linux-based HPC clusters running Slurm or Kubernetes (job scheduling, cgroups, fair-share policies, topology configuration).
  • You possess strong troubleshooting skills in low-level Linux networking, kernel parameters, hardware diagnostics, and storage systems (e.g., NFS, NVMe-oF, or distributed filesystems).
  • You are proficient in automation and infrastructure-as-code tools (e.g., Ansible, Terraform, Helm, Python/Bash scripting).
  • Background supporting distributed deep learning frameworks (DeepSpeed, Ray).

Responsibilities

  • GPU Cluster Orchestration: Design, deploy, and maintain large-scale bare-metal GPU clusters running Slurm and Kubernetes to power multi-node distributed training runs.
  • Workload Management & Scheduler Tuning: Configure, tune, and optimize Slurm schedulers, dynamic job queuing, priority topologies, and autoscaling for optimal GPU utilization and job throughput.
  • Kubernetes Integration: Manage K8s clusters and hybrid Slurm-K8s environments (e.g., KubeRay, MPI Operator, Slurm on K8s) for serving, data processing pipelines, and interactive research environments.
  • Storage & Network Performance: Optimize high-speed interconnects (InfiniBand/RoCE, NCCL, NVLink) and high-throughput distributed storage (Lustre, WEKA, Ceph) to keep tens of thousands of GPUs continuously saturated.
  • Reliability & Monitoring: Build observability pipelines (Prometheus, Grafana, DCGM) to proactively detect hardware degradation, GPU silent errors, network flaps, and node failures before they impact training jobs.

Skills

HPC infrastructure
Linux systems administration
Cluster orchestration
Slurm
Kubernetes
Automation tooling

Education

Bachelor's degree in CS/CE or equivalent in HPC

Tools

KubeRay
MPI Operator
Slurm on Kubernetes
NCCL

Job description

About US

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

Responsibilities
  • GPU Cluster Orchestration: Design, deploy, and maintain large-scale bare-metal GPU clusters running Slurm and Kubernetes to power multi-node distributed training runs.

  • Workload Management & Scheduler Tuning: Configure, tune, and optimize Slurm schedulers, dynamic job queuing, priority topologies, and autoscaling for optimal GPU utilization and job throughput.

  • Kubernetes Integration: Manage K8s clusters and hybrid Slurm-K8s environments (e.g., KubeRay, MPI Operator, Slurm on K8s) for serving, data processing pipelines, and interactive research environments.

  • Storage & Network Performance: Optimize high-speed interconnects (InfiniBand/RoCE, NCCL, NVLink) and high-throughput distributed storage (Lustre, WEKA, Ceph) to keep tens of thousands of GPUs continuously saturated.

  • Reliability & Monitoring: Build observability pipelines (Prometheus, Grafana, DCGM) to proactively detect hardware degradation, GPU silent errors, network flaps, and node failures before they impact training jobs.

Requirements
  • You have a Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in high-performance computing (HPC) or infrastructure engineering.

  • You have deep hands-on expertise administering Linux-based HPC clusters running Slurm or Kubernetes (job scheduling, cgroups, fair-share policies, topology configuration).

  • You possess strong troubleshooting skills in low-level Linux networking, kernel parameters, hardware diagnostics, and storage systems (e.g., NFS, NVMe-oF, or distributed filesystems).

  • You are proficient in automation and infrastructure-as-code tools (e.g., Ansible, Terraform, Helm, Python/Bash scripting).

  • Background supporting distributed deep learning frameworks (DeepSpeed, Ray).

Nice to Have
  • Experience managing high-density GPU infrastructure (NVIDIA H100/B200/B300 systems, DGX/HGX architectures, DCGM, NCCL tuning).

  • Experience with high-speed network fabrics including InfiniBand (Subnet Manager, SMDB) and RoCE (v2).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Toronto

On-site
CAD 120,000 - 165,000
Senior Machine Learning Infrastructure Engineer (Precision, Diagnostics & Hardware)
Senior Machine Learning Infrastructure Engineer (Precision, Diagnostics & Hardware)

Veeda AI • Toronto

On-site
CAD 130,000 - 185,000
Site Reliability Engineer, AI/ML Infrastructure
Site Reliability Engineer, AI/ML Infrastructure

Boson AI • Toronto

On-site
CAD 100,000 - 130,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda AI • Toronto

On-site
CAD 120,000 - 180,000
HPC Specialist
HPC Specialist

Arcadion • Ottawa

On-site
CAD 80,000 - 110,000
HPC Specialist
HPC Specialist

DRW Holdings, LLC. • Montreal (administrative region)

On-site
CAD 100,000 - 130,000
System Software Engineer - AI
System Software Engineer - AI

Delos Data • Montreal (administrative region)

Hybrid
CAD 120,000 - 180,000
Meaningful equity
Benefits
401k
Cluster Operations Software Engineer
Cluster Operations Software Engineer

Cerebras • Toronto

On-site
CAD 120,000 - 190,000
HPC Specialist
HPC Specialist

P2P • Montreal (administrative region)

On-site
CAD 90,000 - 120,000
GPU Cloud Platform Engineer
GPU Cloud Platform Engineer

Yotta Labs • Canada

Remote
CAD 90,000 - 120,000
Flexible work environment
Cutting-edge technology projects
Collaborative team culture