Senior ML Engineer — Training & RL Systems

Socket.dev

Palo Alto (CA)

On-site

USD 195,200 - 262,200

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Career growth and learning
Flexibility and ownership
Collaborative and innovative culture
Impactful AI projects
International environment

Job summary

Nebius is seeking a Senior Machine Learning Engineer to own end-to-end ML work, from strategy to concrete experiments, with hands-on debugging of training and RL pipelines. This role operates at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering.

You will translate ambiguous capability goals into concrete experiments, implement training and RL recipes, and deliver measurable improvements in model quality, throughput, and

Qualifications

  • Strong Python and PyTorch engineering skills.
  • Hands-on experience across two of model training, post-training/RL, applied modeling, data pipelines, or large-scale ML systems.
  • Ability to design rigorous experiments with baselines, ablations, metrics, and failure analysis.
  • Practical understanding of modern LLM behavior, instruction tuning, and evaluation challenges.
  • Practical understanding of transformer training bottlenecks, memory pressure, communication overhead, and checkpointing.
  • Ability to reason quantitatively about model quality, throughput, utilization, reliability, cost, and research velocity.
  • Strong communication skills and ability to collaborate with researchers, engineers, and leadership.

Responsibilities

  • Design and run model-training and post-training experiments, including SFT, continued pretraining, preference optimization (DPO/IPO/KTO), and RL methods such as RLHF/RLAIF, PPO, and GRPO.
  • Build reward functions, judge models, verifiers, task environments, and evaluation sets for reasoning, coding, tool use, and agentic workflows.
  • Create synthetic data and data pipelines, including teacher-student generation, self-play, rejection sampling, filtering, and quality scoring.
  • Analyze model-behavior failures and turn them into targeted data, reward, or algorithm improvements.
  • Build and maintain distributed training and RL infrastructure using Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, Slurm, Kubernetes on large GPU clusters.
  • Implement and debug parallelism strategies (tensor, pipeline, sequence/context, expert, and data parallelism) and build rollout, reward-serving, checkpointing, and experiment-orchestration components.
  • Profile and improve GPU utilization, memory usage, communication efficiency, training throughput, and inference/serving performance.
  • Design rigorous evaluations and ablations for capability, instruction following, reasoning, tool use, safety, and regression risk.
  • Write clear experiment plans, design docs, benchmark reports, and runbooks, and partner across research and platform teams.

Job description

Nebius is seeking a Senior Machine Learning Engineer to own end-to-end ML work, from strategy to concrete experiments, with hands-on debugging of training and RL pipelines. This role operates at the intersection of distributed systems, GPU performance, model training frameworks, RL pipelines, and production engineering.

You will translate ambiguous capability goals into concrete experiments, implement training and RL recipes, and deliver measurable improvements in model quality, throughput, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Research Engineer – AI Agent & RL
Senior ML Research Engineer – AI Agent & RL

Nebius • United States

Remote
USD 180,000 - 280,000
ML Systems Engineer: Distributed GPU Training & RL Pipelines
ML Systems Engineer: Distributed GPU Training & RL Pipelines

Nebius B.V. • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,000 - 263,000
Health insurance
401(k) plan
Parental leave
+2
Lead ML Training Systems Engineer - Multimodal, Large-Scale
Lead ML Training Systems Engineer - Multimodal, Large-Scale

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000
Engineering Manager, ML Infrastructure & Systems
Engineering Manager, ML Infrastructure & Systems

Cursor • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior AI Cloud Infrastructure Engineer
Senior AI Cloud Infrastructure Engineer

Nebius • United States

Remote
USD 150,000 - 190,000
Competitive compensation
Career growth and learning
Flexible ownership and culture
+3
Senior Machine Learning Engineer, Model Training and Reinforcement Learning
Senior Machine Learning Engineer, Model Training and Reinforcement Learning

Socket.dev • Palo Alto (CA)

On-site
USD 195,000 - 263,000
Competitive compensation
Career growth and learning
Flexibility and ownership
+3
Senior Inference & RL Systems Engineer (Scalable ML Infra)
Senior Inference & RL Systems Engineer (Scalable ML Infra)

Magic AI, Inc • San Francisco (CA)

On-site
USD 300,000 - 550,000
Equity compensation
401(k) matching
Health, dental and vision insurance
+4
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Sierracorp • San Francisco (CA)

On-site
USD 150,000 - 200,000
Senior ML Infra Engineer: Scalable AI Training Systems
Senior ML Infra Engineer: Scalable AI Training Systems

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3