ML Systems Engineer — RL & Post-Training Pipelines

Remanence

Paris (TX)

Hybrid

USD 135,000 - 203,000

Full time

12 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Visa sponsorship
Relocation support
Hybrid setup

Job summary

Remanence is pioneering enterprise AI, building intelligent systems that learn continuously from real-world execution. We seek a Member of Technical Staff - Post-training / RL to own post-training methods, training loops, and scalable RL experiments in a hybrid setup with sponsorship for candidates joining us in Paris or London.

You will work on supervised fine-tuning, PPO/GRPO/SDPO, reward computation, and performance optimization across GPUs, with an emphasis on reliability and real-world

Qualifications

  • Experience with large-scale RL and post-training methods may be beneficial.
  • Strong programming and quantitative reasoning.
  • Ability to move from experiments to profiler traces.

Responsibilities

  • Implement and improve supervised fine-tuning, preference optimization, and RL methods such as PPO, GRPO and SDPO.
  • Own the training loop from rollout generation and reward computation through policy updates, weight synchronization, and checkpointing.
  • Improve throughput, GPU utilization, and memory efficiency through batching, parallelism, communication overlap, and asynchronous execution.
  • Profile bottlenecks and investigate numerical instability, stale-policy effects, and training-inference discrepancies. Verify that optimizations preserve learning behavior.
  • Occasionally deploy models in client environments and hill-climb alongside their teams: analyze real failures, refine rewards and training data, and iterate on task success, latency, and inference cost.

Tools

PyTorch
Miles
NVIDIA Molt
verl
FSDP
Megatron-LM
vLLM
SGLang
Ray
NVIDIA Nsight Systems
CUDA
Triton
CuTE
PyTorch Profiler

Job description

Remanence is pioneering enterprise AI, building intelligent systems that learn continuously from real-world execution. We seek a Member of Technical Staff - Post-training / RL to own post-training methods, training loops, and scalable RL experiments in a hybrid setup with sponsorship for candidates joining us in Paris or London.

You will work on supervised fine-tuning, PPO/GRPO/SDPO, reward computation, and performance optimization across GPUs, with an emphasis on reliability and real-world

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff — Post-training / RL
Member of Technical Staff — Post-training / RL

Remanence • Paris (TX)

Hybrid
USD 135,000 - 203,000
Visa sponsorship
Relocation support
Hybrid setup
ML Systems Engineer: Post-Training Pipelines & RL
ML Systems Engineer: Post-Training Pipelines & RL

ThirdLayer, Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
ML Systems Engineer — RL Training & Finetuning
ML Systems Engineer — RL Training & Finetuning

Anthropic • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Competitive compensation
Equity donation matching
Generous vacation and parental leave
+1
Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
RL Post-Training Systems Architect (Equity Eligible)
RL Post-Training Systems Architect (Equity Eligible)

NVIDIA Gruppe • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Staff Engineer - AI Systems & Post-Training RL
Staff Engineer - AI Systems & Post-Training RL

Cohere • New York (NY)

On-site
USD 150,000 - 230,000
Lunch stipend and health benefits
Dental benefits
RRSP/401K/Pension
+6
ML Systems Engineer: RL Training & AI Safety
ML Systems Engineer: RL Training & AI Safety

Anthropic • Seattle (WA)

Hybrid
USD 520,000 - 850,000
Equity donation matching
Generous vacation
Parental leave
+2
ML Infra Engineer - Scale RL Pipelines (Paris)
ML Infra Engineer - Scale RL Pipelines (Paris)

United States Digital Space LLC • Paris (TX)

Hybrid
USD 125,000 - 182,000
Competitive salary + equity
Hybrid work from Paris with relocation
Top-tier medical insurance in France
+2
Senior RL Post-Training Systems Engineer (Equity)
Senior RL Post-Training Systems Engineer (Equity)

NVIDIA • Washington

On-site
USD 184,000 - 357,000
Equity
Comprehensive benefits
Lead ML Training Systems Engineer - Multimodal, Large-Scale
Lead ML Training Systems Engineer - Multimodal, Large-Scale

Rhoda AI • Palo Alto (CA)

On-site
USD 210,000 - 320,000