Lead ML Systems Engineer - Large-Scale Training & RL Infra

Nebius

Palo Alto (CA)

Remote

USD 195,200 - 262,200

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) plan
Parental leave
Remote work reimbursement
Disability & life insurance

Job summary

Nebius is building a world-class AI training and model post-training capability. This Senior ML Systems Engineer role owns end-to-end training infrastructure for large-scale experiments, focusing on reliability, reproducibility, and efficiency.

You will be hands-on, debug complex distributed training failures, and push measurable improvements in throughput, stability, and GPU utilization while collaborating with researchers and platform engineers.

Qualifications

  • Strong Python and PyTorch engineering skills.
  • Hands-on experience with distributed model training, large-scale ML systems, or GPU cluster workloads.
  • Practical understanding of transformer training bottlenecks, memory pressure, gradient/optimizer state, communication overhead, and checkpointing.
  • Experience debugging production or research training jobs across multiple GPUs or nodes.
  • Ability to reason quantitatively about throughput, utilization, memory, reliability, cost, and research velocity.
  • Strong communication skills and ability to collaborate with researchers, ML engineers, platform engineers, and leadership.
  • Nice-to-have: Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, Slurm, Kubernetes, or large internal training platforms.
  • Nice-to-have: RL infrastructure frameworks such as verl, slime, AReaL, OpenRLHF, TRL, or PPO/GRPO/RLHF systems.
  • Nice-to-have: Familiarity with NCCL, CUDA, Triton, InfiniBand, RDMA, RoCE, H100/H200/B200 clusters.

Responsibilities

  • Build and maintain distributed training infrastructure for SFT, continued pretraining, preference optimization, and RL workloads.
  • Integrate and extend frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF, or equivalent internal systems.
  • Implement and debug parallelism strategies including tensor, pipeline, sequence/context, expert, and data parallelism.
  • Build reliable rollout, reward model serving, replay/data buffer, checkpointing, evaluation, and experiment orchestration components for RL training.
  • Profile and improve GPU utilization, communication efficiency, memory usage, and training throughput.
  • Diagnose failures across NCCL, CUDA, PyTorch, Ray, schedulers, storage, networking, and checkpointing layers.
  • Create reproducible training runs, launch scripts, dashboards, runbooks, and operating guides.
  • Partner with research scientists to turn algorithmic training recipes into scalable, debuggable systems.
  • Write clear design docs, incident reports, benchmark reports, and operating guides.

Skills

Python
PyTorch
Distributed training
GPU clusters
Profiling

Tools

Megatron-LM
DeepSpeed
PyTorch FSDP/DTensor
Ray
Slurm
Kubernetes

Job description

Nebius is building a world-class AI training and model post-training capability. This Senior ML Systems Engineer role owns end-to-end training infrastructure for large-scale experiments, focusing on reliability, reproducibility, and efficiency.

You will be hands-on, debug complex distributed training failures, and push measurable improvements in throughput, stability, and GPU utilization while collaborating with researchers and platform engineers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Engineer — Training & RL Systems
Senior ML Engineer — Training & RL Systems

Socket.dev • Palo Alto (CA)

On-site
USD 195,000 - 263,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
ML Systems Engineer: Distributed GPU Training & RL Pipelines
ML Systems Engineer: Distributed GPU Training & RL Pipelines

Nebius B.V. • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior ML Research Engineer – AI Agent & RL
Senior ML Research Engineer – AI Agent & RL

Nebius • United States

Remote
USD 180,000 - 280,000
Senior AI Cloud Infrastructure Engineer
Senior AI Cloud Infrastructure Engineer

Nebius • United States

Remote
USD 150,000 - 190,000
Competitive compensation
Career growth and learning
Flexible ownership and culture
+3
Lead ML Training Systems Engineer - Multimodal, Large-Scale
Lead ML Training Systems Engineer - Multimodal, Large-Scale

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Senior ML Infra Engineer: Scalable AI Training Systems
Senior ML Infra Engineer: Scalable AI Training Systems

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,000 - 263,000
Health insurance
401(k) plan
Parental leave
+2
Engineering Manager, ML Infrastructure & Systems
Engineering Manager, ML Infrastructure & Systems

Cursor • San Francisco (CA)

On-site
USD 180,000 - 260,000
Lead ML Systems Engineer — Distributed GPU Training & Infra
Lead ML Systems Engineer — Distributed GPU Training & Infra

Nvidia Corporation • Santa Clara (CA)

On-site
USD 224,000 - 431,000
Equity
Benefits
ML Systems Engineer, Large-Scale Model Training & RL Infrastructure
ML Systems Engineer, Large-Scale Model Training & RL Infrastructure

Nebius • Palo Alto (CA)

On-site
USD 195,200 - 262,200
Health insurance
401(k) plan
Parental leave
+2