Member of Technical Staff - ML Performance

Veeda AI

Toronto

On-site

CAD 120,000 - 180,000

Full time

9 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Veeda AI in Toronto is seeking a Member of Technical Staff - ML Performance to own step time and model throughput for multi-node video world model training, optimize tensor usage, and implement advanced parallelism.

You will write and tune CUDA and Triton kernels, improve precision and numerical stability, and build fault diagnostics and elastic checkpointing to minimize downtime during interruptions.

Qualifications

  • Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related field.
  • Hands-on PyTorch experience with large-scale parallelism stacks on multi-node jobs.
  • Fluency in Python and C++/CUDA with the ability to predict where a kernel will stall from memory access pattern before profiling.
  • Experience profiling live training runs with Nsight Systems or the PyTorch profiler and translating traces into quantifiable step-time or MFU improvements.
  • Expertise in at least one of low-precision numerics, kernel authoring, or large-run fault diagnosis.

Responsibilities

  • Distributed Training Throughput: Own step time and model FLOPs utilization for multi-node video world model training, choosing tensor and expert parallelism.
  • Precision & Numerical Stability: Take BF16, FP8, and NVFP4 recipes from running to converging on Blackwell, addressing scaling-factor and accumulation bugs in video tokenizer and VAE layers.
  • Kernels & Compilation: Write and tune CUDA and Triton kernels to enable FlashAttention-4, FlexAttention, and torch.compile for long sequences.
  • Communication & Overlap: Tune NCCL collectives and compute/communication overlap across NVLink domains and the fabric.
  • Fault Diagnostics & Recovery: Build detection for silent data corruption, stuck kernels, and asynchronous checkpointing to minimize interruption cost.

Skills

Python
C++/CUDA
PyTorch
Profiling

Education

Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field

Tools

Nsight Systems
PyTorch Profiler
Megatron-Core
DeepSpeed
TorchTitan
CUDA
Triton

Job description

Member of Technical Staff - ML Performance
About Us

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

Responsibilities
  • Distributed Training Throughput: Own step time and model FLOPs utilization for multi-node video world model training, choosing the tensor, context, and expert parallelism mix in PyTorch FSDP2 and Megatron-Core rather than inheriting a default.

  • Precision & Numerical Stability: Take BF16, FP8, and NVFP4 recipes from running to converging on Blackwell, chasing scaling-factor and accumulation bugs into the video tokenizer and VAE layers where the activation outliers actually live.

  • Kernels & Compilation: Write and tune the CUDA and Triton kernels PyTorch does not give us, driving FlashAttention-4, FlexAttention, and torch.compile integration so quadratic attention over long video sequences stops setting step time.

  • Communication & Overlap: Tune NCCL collectives and compute/communication overlap across NVLink domains and the fabric, using the NCCL flight recorder to turn a watchdog timeout into a named rank and collective, not a restart.

  • Fault Diagnostics & Recovery: Build the detection layer for silent data corruption (SDC), stuck CUDA kernels, and "card-freeze" hangs, plus asynchronous and tiered checkpointing that makes an interruption cost minutes rather than a day.

Requirements
  • Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field.

  • Deep hands-on experience with PyTorch and at least one large-scale parallelism stack (FSDP2, Megatron-Core, TorchTitan, or DeepSpeed) on real multi-node jobs, not single-node approximations.

  • Fluency in Python and C++/CUDA with the ability to predict where a kernel will stall from its memory access pattern before profiling.

  • Experience profiling live training runs with Nsight Systems or the PyTorch profiler and translating traces into quantifiable step-time or MFU improvements.

  • Expertise in at least one of low-precision numerics, kernel authoring, or large-run fault diagnosis, and credibility in the others.

Nice to Have
  • Experience writing Triton, CUTLASS, or CuTe-DSL kernels, or contributing to open-source kernel libraries.

  • Experience implementing context or sequence parallelism for long-horizon video or high-token-count models.

  • Experience running or porting large training workloads on AMD GPUs (ROCm) or Google TPUs (JAX/XLA).

  • Experience building fault-tolerant training with elastic world size, dynamic node re-queueing, or asynchronous distributed checkpointing.

  • Experience optimizing generative inference for interactive rollouts, including few-step samplers, distillation, and KV or latent caching.

  • Publications or presentations on machine learning systems, compilers, or high-performance kernels.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - World Models
Member of Technical Staff - World Models

Veeda Innovation • Toronto

On-site
CAD 120,000 - 180,000
Member of Technical Staff - ML Operations
Member of Technical Staff - ML Operations

Veeda AI • Toronto

On-site
CAD 120,000 - 150,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Toronto

On-site
CAD 120,000 - 165,000
Member of Technical Staff - Data
Member of Technical Staff - Data

Veeda AI • Toronto

On-site
CAD 120,000 - 180,000
Member of Technical Staff - Simulation
Member of Technical Staff - Simulation

Veeda AI • Toronto

On-site
CAD 120,000 - 180,000
Senior ML Performance Engineer - Distributed Training
Senior ML Performance Engineer - Distributed Training

Veeda AI • Toronto

On-site
CAD 120,000 - 180,000
Member of Technical Staff - Robotics
Member of Technical Staff - Robotics

Veeda AI • Toronto

On-site
CAD 110,000 - 150,000
Member of Technical Staff - Robotics
Member of Technical Staff - Robotics

Veeda Innovation • Toronto

On-site
CAD 90,000 - 160,000
Software Engineer- Model Performance Systems
Software Engineer- Model Performance Systems

Baseten • Montreal (administrative region)

On-site
CAD 225,067 - 281,334
Equity
Medical, dental, vision insurance
Flexible PTO
+4
Inference Performance Engineer
Inference Performance Engineer

adaption • Quebec

On-site
CAD 110,000 - 170,000
Flexible work
Adaption Passport
Lunch Stipend
+1