Get more replies from employers
Send a job-specific resume in minutes.
Veeda AI is seeking a Member of Technical Staff - ML Performance to optimize distributed training for multi-node video world models. You will own step time, FLOPs utilization, memory patterns, and parallelism strategies across NVLink, NCCL, and PyTorch FSDP2 in real workloads.
You’ll implement kernels, tune CUDA/Triton components, improve precision stability, and contribute to fault diagnostics and elastic checkpointing.
Veeda AI is seeking a Member of Technical Staff - ML Performance to optimize distributed training for multi-node video world models. You will own step time, FLOPs utilization, memory patterns, and parallelism strategies across NVLink, NCCL, and PyTorch FSDP2 in real workloads.
You’ll implement kernels, tune CUDA/Triton components, improve precision stability, and contribute to fault diagnostics and elastic checkpointing.