ML Systems Engineer: Distributed GPU Training & RL Pipelines

Nebius B.V.

Palo Alto (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Nebius B.V. is building an AI training and model post-training capability focused on frontier models. This role owns the infrastructure for large-scale training, RL experiments, and production-grade workflows.

You will work at the intersection of distributed systems, GPU performance, and ML framework integration. The role requires strong Python and PyTorch engineering skills, hands-on experience with distributed model training, and the ability to optimize throughput and memory across multi-GPU

Qualifications

  • Proven Python and PyTorch engineering experience.
  • Hands-on distributed training of large-scale ML systems.
  • Understanding of transformer bottlenecks and memory pressure.
  • Experience debugging training across multiple GPUs/nodes.

Responsibilities

  • Build and maintain distributed training infrastructure for SFT, continued pretraining, preference optimization, and RL workloads.
  • Integrate frameworks like Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, and other internal systems.
  • Implement and debug parallelism: tensor, pipeline, sequence/context, expert, data parallelism.
  • Build rollout, reward model serving, replay/data buffer, checkpointing, evaluation, and orchestration components for RL training.
  • Profile GPU utilization, memory, and throughput; optimize communication.
  • Diagnose failures across NCCL, CUDA, PyTorch, Ray, schedulers, storage, networking, and checkpointing layers.
  • Create reproducible training runs, launch scripts, dashboards, runbooks, and operat­ing tooling for researchers.
  • Collaborate with researchers to turn algorithms into scalable, debuggable systems.
  • Write design docs, incident reports, benchmarks, and operating guides.

Skills

Python programming
PyTorch
Distributed systems
GPU clustering
Performance optimization
Experimentation & benchmarking
Debugging complex ML workloads
Communication & collaboration

Tools

Megatron-LM
DeepSpeed
PyTorch FSDP/DTensor
Ray
Slurm
Kubernetes

Job description

Nebius B.V. is building an AI training and model post-training capability focused on frontier models. This role owns the infrastructure for large-scale training, RL experiments, and production-grade workflows.

You will work at the intersection of distributed systems, GPU performance, and ML framework integration. The role requires strong Python and PyTorch engineering skills, hands-on experience with distributed model training, and the ability to optimize throughput and memory across multi-GPU

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer — Training & RL Systems
Senior ML Engineer — Training & RL Systems

Socket.dev • Palo Alto (CA)

On-site
USD 195,000 - 263,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
ML Systems Engineer
ML Systems Engineer

Nebius B.V. • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Lead ML Systems Engineer — Distributed GPU Training & Infra
Lead ML Systems Engineer — Distributed GPU Training & Infra

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
Staff ML Infra Engineer: Distributed Training & Inference
Staff ML Infra Engineer: Distributed Training & Inference

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000
Distributed AI Training Architect
Distributed AI Training Architect

cerence • United States

On-site
USD 180,000 - 250,000
Senior RL Post-Training Systems Engineer
Senior RL Post-Training Systems Engineer

NVIDIA • Indiana (PA)

On-site
USD 150,000 - 190,000
Senior RL Post-Training Systems Engineer
Senior RL Post-Training Systems Engineer

NVIDIA • Massachusetts

On-site
USD 224,000 - 357,000
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • Indiana (PA)

On-site
USD 150,000 - 190,000
Senior ML Engineer - GPU Inference & Large-Scale AI Cloud
Senior ML Engineer - GPU Inference & Large-Scale AI Cloud

Nebius • United States

Remote
USD 180,000 - 250,000
Senior AI Systems Architect: Scalable Distributed Training
Senior AI Systems Architect: Scalable Distributed Training

Cerence AI • United States

On-site
USD 180,000 - 240,000