ML Performance Engineer — Distributed Training & Kernels

Veeda

California (MO)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Veeda AI is seeking a Member of Technical Staff - ML Performance to optimize distributed training for multi-node video world models. You will own step time, FLOPs utilization, memory patterns, and parallelism strategies across NVLink, NCCL, and PyTorch FSDP2 in real workloads.

You’ll implement kernels, tune CUDA/Triton components, improve precision stability, and contribute to fault diagnostics and elastic checkpointing.

Qualifications

  • Bachelor's degree or equivalent hands-on experience in Computer Science, Computer Engineering, or a related technical field.
  • Deep hands-on experience with PyTorch and at least one large-scale parallelism stack on real multi-node jobs.
  • Fluency in Python and C++/CUDA with the ability to predict where a kernel will stall from its memory access pattern before profiling.
  • Experience profiling live training runs with Nsight Systems or the PyTorch profiler and translating traces into quantifiable step-time or MFU improvements.
  • Expertise in at least one of low-precision numerics, kernel authoring, or large-run fault diagnosis, and credibility in the others.

Responsibilities

  • Own step time and model FLOPs utilization for multi-node video world model training, choosing the tensor, context, and expert parallelism mix in PyTorch FSDP2 and Megatron-Core rather than inheriting a default.
  • Take BF16, FP8, and NVFP4 recipes from running to converging on Blackwell, chasing scaling-factor and accumulation bugs into the video tokenizer and VAE layers where the activation outliers actually live.
  • Write and tune the CUDA and Triton kernels PyTorch does not give us, driving FlashAttention-4, FlexAttention, and torch.compile integration so quadratic attention over long video sequences stops setting step time.
  • Tune NCCL collectives and compute/communication overlap across NVLink domains and the fabric, using the NCCL flight recorder to turn a watchdog timeout into a named rank and collective, not a restart.
  • Build the detection layer for silent data corruption (SDC), stuck CUDA kernels, and "card-freeze" hangs, plus asynchronous and tiered checkpointing that makes an interruption cost minutes rather than a day.

Skills

PyTorch
Python
C++/CUDA
Multi-node training
Nsight Systems
PyTorch profiler
FSDP2
Megatron-Core
DeepSpeed
Kernels tuning

Education

Bachelor's degree in CS/CE or related field

Tools

Triton
CUTLASS
CuTe-DSL
ROCm
JAX/XLA

Job description

Veeda AI is seeking a Member of Technical Staff - ML Performance to optimize distributed training for multi-node video world models. You will own step time, FLOPs utilization, memory patterns, and parallelism strategies across NVLink, NCCL, and PyTorch FSDP2 in real workloads.

You’ll implement kernels, tune CUDA/Triton components, improve precision stability, and contribute to fault diagnostics and elastic checkpointing.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Engineer — Performance & Scale
ML Systems Engineer — Performance & Scale

Veeda Innovation • Northern (KY)

Hybrid
USD 150,000 - 230,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda • California (MO)

On-site
USD 180,000 - 240,000
Member of Technical Staff - ML Performance
Member of Technical Staff - ML Performance

Veeda Innovation • Northern (KY)

Hybrid
USD 150,000 - 230,000
Performance Engineer - ML Training & CUDA Kernels
Performance Engineer - ML Training & CUDA Kernels

Cohere • New York (NY)

Hybrid
USD 150,000 - 190,000
Lunch stipend
Health benefits
RRSP matching
+5
Senior AI Training Performance Engineer (GPU & Scale)
Senior AI Training Performance Engineer (GPU & Scale)

figure.ai • San Jose (CA), Northern (KY)

Hybrid
USD 200,000 - 400,000
Distributed ML Training Performance Engineer
Distributed ML Training Performance Engineer

OpenAI • California (MO)

Hybrid
USD 170,000 - 260,000
Relocation assistance
AI Infrastructure Architect for HPC & GPU Clusters
AI Infrastructure Architect for HPC & GPU Clusters

Veeda • California (MO)

On-site
USD 150,000 - 200,000
Kernel Engineer for High-Performance ML Compute (CUDA/Triton)
Kernel Engineer for High-Performance ML Compute (CUDA/Triton)

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
AI Infrastructure Engineer: HPC GPU Clusters
AI Infrastructure Engineer: HPC GPU Clusters

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior AI Inference Performance Engineer
Senior AI Inference Performance Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000