Senior ML Performance Engineer - Distributed Training

Veeda AI

Toronto

On-site

CAD 120,000 - 180,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Veeda AI in Toronto is seeking a Member of Technical Staff - ML Performance to own step time and model throughput for multi-node video world model training, optimize tensor usage, and implement advanced parallelism.

You will write and tune CUDA and Triton kernels, improve precision and numerical stability, and build fault diagnostics and elastic checkpointing to minimize downtime during interruptions.

Qualifications

  • Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related field.
  • Hands-on PyTorch experience with large-scale parallelism stacks on multi-node jobs.
  • Fluency in Python and C++/CUDA with the ability to predict where a kernel will stall from memory access pattern before profiling.
  • Experience profiling live training runs with Nsight Systems or the PyTorch profiler and translating traces into quantifiable step-time or MFU improvements.
  • Expertise in at least one of low-precision numerics, kernel authoring, or large-run fault diagnosis.

Responsibilities

  • Distributed Training Throughput: Own step time and model FLOPs utilization for multi-node video world model training, choosing tensor and expert parallelism.
  • Precision & Numerical Stability: Take BF16, FP8, and NVFP4 recipes from running to converging on Blackwell, addressing scaling-factor and accumulation bugs in video tokenizer and VAE layers.
  • Kernels & Compilation: Write and tune CUDA and Triton kernels to enable FlashAttention-4, FlexAttention, and torch.compile for long sequences.
  • Communication & Overlap: Tune NCCL collectives and compute/communication overlap across NVLink domains and the fabric.
  • Fault Diagnostics & Recovery: Build detection for silent data corruption, stuck kernels, and asynchronous checkpointing to minimize interruption cost.

Skills

Python
C++/CUDA
PyTorch
Profiling

Education

Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field

Tools

Nsight Systems
PyTorch Profiler
Megatron-Core
DeepSpeed
TorchTitan
CUDA
Triton

Job description

Veeda AI in Toronto is seeking a Member of Technical Staff - ML Performance to own step time and model throughput for multi-node video world model training, optimize tensor usage, and implement advanced parallelism.

You will write and tune CUDA and Triton kernels, improve precision and numerical stability, and build fault diagnostics and elastic checkpointing to minimize downtime during interruptions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda AI • Toronto

On-site
CAD 120,000 - 180,000
Senior ML Infra Engineer: Scale GPU Training & Reliability
Senior ML Infra Engineer: Scale GPU Training & Reliability

Veeda AI • Toronto

On-site
CAD 130,000 - 185,000
Senior Machine Learning Infrastructure Engineer (Precision, Diagnostics & Hardware)
Senior Machine Learning Infrastructure Engineer (Precision, Diagnostics & Hardware)

Veeda AI • Toronto

On-site
CAD 130,000 - 185,000
ML Kernel Performance Eng. Manager, Accelerators
ML Kernel Performance Eng. Manager, Accelerators

Socket.dev • Toronto

On-site
CAD 171,000 - 286,000
ML Ops Engineer - Reproducible AI Pipelines
ML Ops Engineer - Reproducible AI Pipelines

Veeda AI • Toronto

On-site
CAD 120,000 - 150,000
Senior AI Infrastructure Engineer — HPC & Compute Clusters
Senior AI Infrastructure Engineer — HPC & Compute Clusters

Veeda AI • Toronto

On-site
CAD 100,000 - 150,000
Senior AI Inference Systems Engineer
Senior AI Inference Systems Engineer

NVIDIA Corporation • Toronto

Hybrid
CAD 170,000 - 275,000
Equity
Benefits
Senior Production AI Systems Engineer
Senior Production AI Systems Engineer

Jaide Health • Toronto

Hybrid
CAD 140,000 - 210,000
Lunch stipend
Health and dental benefits
RRSP matching
+5
Member of Technical Staff - ML Operations
Member of Technical Staff - ML Operations

Veeda AI • Toronto

On-site
CAD 120,000 - 150,000
DL Performance Software Engineer - LLM Inference
DL Performance Software Engineer - LLM Inference

NVIDIA Gruppe • Toronto

Hybrid
CAD 135,000 - 220,000
Equity
Benefits
Hybrid work model