Senior ML Infra Engineer: Scale GPU Training & Reliability

Veeda AI

Toronto

On-site

CAD 130,000 - 185,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We seek engineers to design and optimize distributed training systems across large GPU clusters, handling FP16/BF16/FP8 precision and debugging complex stability issues.

You will implement fault-detection, performance profiling, and resilient checkpointing to keep researchers productive in a fast-moving environment. Collaboration across teams is essential for success.

Qualifications

  • Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field.
  • Experience with deep learning training frameworks (e.g., PyTorch) and distributed training paradigms (FSDP, Megatron-LM, DeepSpeed, Tensor Parallelism, Pipeline Parallelism).
  • Proven experience in numerical precision analysis, low-precision training (BF16/FP8), and debugging complex loss divergence/stability issues in massive training runs.
  • Strong root-cause analysis skills for hardware/software interaction bugs, including stuck CUDA kernels, NCCL timeouts, GPU hardware faults, and silent training corruptions.

Responsibilities

  • Design, optimize, and maintain high-throughput distributed training systems across large-scale GPU clusters for multi-modal foundation models.
  • Debug, diagnose, and resolve subtle numerical instability issues (underflow/overflow, loss spikes, gradient explosion, and mixed-precision divergence) in FP16, BF16, FP8, and custom quantization schemes.
  • Build advanced fault-detection mechanisms and automated diagnostics to rapidly pinpoint and isolate silent data corruption (SDC), hardware hang/deadlock, memory leaks, and card-freeze issues during large training runs.
  • Profile distributed communication bottlenecks, memory usage, and kernel execution to improve overall FLOPS utilization across multi-node, multi-GPU training jobs.
  • Develop resilient checkpointing systems, rapid fault-recovery pipelines, and execution telemetry to keep researcher productivity high and hardware downtime minimal.

Skills

Python
C++/CUDA
PyTorch
Distributed training
Megatron-LM
DeepSpeed
Tensor Parallelism
Pipeline Parallelism
NCCL
Low-level GPU
CUDA kernels

Education

Bachelor's degree in CS/CE

Tools

Triton kernels
NCCL
DeepSpeed
Megatron-LM
PyTorch

Job description

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We seek engineers to design and optimize distributed training systems across large GPU clusters, handling FP16/BF16/FP8 precision and debugging complex stability issues.

You will implement fault-detection, performance profiling, and resilient checkpointing to keep researchers productive in a fast-moving environment. Collaboration across teams is essential for success.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Machine Learning Infrastructure Engineer (Precision, Diagnostics & Hardware)
Senior Machine Learning Infrastructure Engineer (Precision, Diagnostics & Hardware)

Veeda AI • Toronto

On-site
CAD 130,000 - 185,000
Senior ML Performance Engineer - Distributed Training
Senior ML Performance Engineer - Distributed Training

Veeda AI • Toronto

On-site
CAD 120,000 - 180,000
Senior AI Infrastructure Engineer — HPC & Compute Clusters
Senior AI Infrastructure Engineer — HPC & Compute Clusters

Veeda AI • Toronto

On-site
CAD 100,000 - 150,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda AI • Toronto

On-site
CAD 120,000 - 180,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Toronto

On-site
CAD 120,000 - 165,000
Senior AI Inference Systems Engineer
Senior AI Inference Systems Engineer

NVIDIA • Toronto

On-site
CAD 170,000 - 275,000
Equity
Benefits
Member of Technical Staff - ML Operations
Member of Technical Staff - ML Operations

Veeda AI • Toronto

On-site
CAD 120,000 - 150,000
Senior Software Engineer, AI Inference Systems
Senior Software Engineer, AI Inference Systems

NVIDIA • Toronto

On-site
CAD 170,000 - 275,000
Equity
Benefits
Staff ML Systems Engineer - Model Efficiency & Inference
Staff ML Systems Engineer - Model Efficiency & Inference

Cohere • Toronto

Hybrid
CAD 140,000 - 200,000
Lunch stipend
Health benefits
Retirement plan matching (RRSP/401K)
Senior Software Engineer, AI Inference Systems
Senior Software Engineer, AI Inference Systems

NVIDIA Corporation • Toronto

On-site
CAD 170,000 - 275,000
Equity
Benefits