Get more replies from employers
Send a job-specific resume in minutes.
Nebius B.V. is building an AI training and model post-training capability focused on frontier models. This role owns the infrastructure for large-scale training, RL experiments, and production-grade workflows.
You will work at the intersection of distributed systems, GPU performance, and ML framework integration. The role requires strong Python and PyTorch engineering skills, hands-on experience with distributed model training, and the ability to optimize throughput and memory across multi-GPU
Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering.