Get more replies from employers
Send a job-specific resume in minutes.
Veeda AI is seeking a Member of Technical Staff - ML Performance to optimize distributed training and inference for large-scale video models. You will drive throughput, profile bottlenecks, and implement kernel-level enhancements across PyTorch, CUDA, and specialized parallelism stacks.
The role requires deep experience with FSDP2, Megatron-Core, TorchTitan, or DeepSpeed, plus strong Python/C++ skills. Location in the US is expected, with a focus on scalable acceleration across GPU clusters.
Veeda AI is seeking a Member of Technical Staff - ML Performance to optimize distributed training and inference for large-scale video models. You will drive throughput, profile bottlenecks, and implement kernel-level enhancements across PyTorch, CUDA, and specialized parallelism stacks.
The role requires deep experience with FSDP2, Megatron-Core, TorchTitan, or DeepSpeed, plus strong Python/C++ skills. Location in the US is expected, with a focus on scalable acceleration across GPU clusters.