ML Systems Engineer

Nebius B.V.

Palo Alto (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nebius B.V. is building an AI training and model post-training capability focused on frontier models. This role owns the infrastructure for large-scale training, RL experiments, and production-grade workflows.

You will work at the intersection of distributed systems, GPU performance, and ML framework integration. The role requires strong Python and PyTorch engineering skills, hands-on experience with distributed model training, and the ability to optimize throughput and memory across multi-GPU

Qualifications

  • Proven Python and PyTorch engineering experience.
  • Hands-on distributed training of large-scale ML systems.
  • Understanding of transformer bottlenecks and memory pressure.
  • Experience debugging training across multiple GPUs/nodes.

Responsibilities

  • Build and maintain distributed training infrastructure for SFT, continued pretraining, preference optimization, and RL workloads.
  • Integrate frameworks like Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, and other internal systems.
  • Implement and debug parallelism: tensor, pipeline, sequence/context, expert, data parallelism.
  • Build rollout, reward model serving, replay/data buffer, checkpointing, evaluation, and orchestration components for RL training.
  • Profile GPU utilization, memory, and throughput; optimize communication.
  • Diagnose failures across NCCL, CUDA, PyTorch, Ray, schedulers, storage, networking, and checkpointing layers.
  • Create reproducible training runs, launch scripts, dashboards, runbooks, and operat­ing tooling for researchers.
  • Collaborate with researchers to turn algorithms into scalable, debuggable systems.
  • Write design docs, incident reports, benchmarks, and operating guides.

Skills

Python programming
PyTorch
Distributed systems
GPU clustering
Performance optimization
Experimentation & benchmarking
Debugging complex ML workloads
Communication & collaboration

Tools

Megatron-LM
DeepSpeed
PyTorch FSDP/DTensor
Ray
Slurm
Kubernetes

Job description

Nebius Token Factory is building an AI training and model post-training capability for frontier model improvement. This role owns the infrastructure that makes large-scale training and RL experiments possible, reliable, reproducible, and efficient. The work sits at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering.

Your responsibilities
  • Build and maintain distributed training infrastructure for SFT, continued pretraining, preference optimization, and RL workloads.
  • Integrate and extend frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF, or equivalent internal systems.
  • Implement and debug parallelism strategies including tensor, pipeline, sequence/context, expert, and data parallelism.
  • Build reliable rollout, reward model serving, replay/data buffer, checkpointing, evaluation, and experiment orchestration components for RL training.
  • Profile and improve GPU utilization, communication efficiency, memory usage, and training throughput.
  • Diagnose failures across NCCL, CUDA, PyTorch, Ray, schedulers, storage, networking, and checkpointing layers.
  • Create reproducible training runs, launch scripts, dashboards, runbooks, and operational tooling for research users.
  • Partner with research scientists to turn algorithmic training recipes into scalable, debuggable systems.
  • Write clear design docs, incident reports, benchmark reports, and operating guides.
Must-haves
  • Strong Python and PyTorch engineering skills.
  • Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads.
  • Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing.
  • Experience debugging production or research training jobs across multiple GPUs or nodes.
  • Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity.
  • Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership.
Nice-to-haves
  • Experience with Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, Slurm, Kubernetes, or large internal training platforms.
  • Experience with RL infrastructure frameworks such as verl, slime, AReaL, OpenRLHF, TRL, or custom PPO/GRPO/RLHF systems.
  • Familiarity with NCCL, CUDA, Triton, Nsight, InfiniBand, RDMA, RoCE, H100/H200/B200 clusters, or storage/network bottlenecks.
  • Experience supporting SFT, DPO, PPO, GRPO, RLAIF, reward model serving, rollout generation, or agent training workloads.
  • Open-source contributions to distributed training, RL infrastructure, PyTorch, Ray, Megatron, DeepSpeed, or related systems. Strong Python and PyTorch engineering skills, Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads, Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing, Experience debugging production or research training jobs across multiple GPUs or nodes, Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity, Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • Indiana (PA)

On-site
USD 150,000 - 190,000
Research Engineer
Research Engineer

Mind Robotics • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, ML Engineer
Member of Technical Staff, ML Engineer

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000
Research Member of Technical Staff- Training Systems
Research Member of Technical Staff- Training Systems

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Member of Technical Staff — Training Infrastructure
Member of Technical Staff — Training Infrastructure

Kindredventures • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Training Infra
Member of Technical Staff, Training Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Systems Engineer: Distributed GPU Training & RL Pipelines
ML Systems Engineer: Distributed GPU Training & RL Pipelines

Nebius B.V. • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior Principal AI Engineer
Senior Principal AI Engineer

cerence • United States

On-site
USD 180,000 - 250,000
Reinforcement Learning Infrastructure Engineer
Reinforcement Learning Infrastructure Engineer

Jobtailor • Palo Alto (CA)

On-site
USD 150,000 - 210,000