Member of Technical Staff — Post-training / RL

Remanence

Paris (TX)

Hybrid

USD 135,000 - 203,000

Full time

12 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Visa sponsorship
Relocation support
Hybrid setup

Job summary

Remanence is pioneering enterprise AI, building intelligent systems that learn continuously from real-world execution. We seek a Member of Technical Staff - Post-training / RL to own post-training methods, training loops, and scalable RL experiments in a hybrid setup with sponsorship for candidates joining us in Paris or London.

You will work on supervised fine-tuning, PPO/GRPO/SDPO, reward computation, and performance optimization across GPUs, with an emphasis on reliability and real-world

Qualifications

  • Experience with large-scale RL and post-training methods may be beneficial.
  • Strong programming and quantitative reasoning.
  • Ability to move from experiments to profiler traces.

Responsibilities

  • Implement and improve supervised fine-tuning, preference optimization, and RL methods such as PPO, GRPO and SDPO.
  • Own the training loop from rollout generation and reward computation through policy updates, weight synchronization, and checkpointing.
  • Improve throughput, GPU utilization, and memory efficiency through batching, parallelism, communication overlap, and asynchronous execution.
  • Profile bottlenecks and investigate numerical instability, stale-policy effects, and training-inference discrepancies. Verify that optimizations preserve learning behavior.
  • Occasionally deploy models in client environments and hill-climb alongside their teams: analyze real failures, refine rewards and training data, and iterate on task success, latency, and inference cost.

Tools

PyTorch
Miles
NVIDIA Molt
verl
FSDP
Megatron-LM
vLLM
SGLang
Ray
NVIDIA Nsight Systems
CUDA
Triton
CuTE
PyTorch Profiler

Job description

Member of Technical Staff - Post-training / RL

Remanence is pioneering the next era of enterprise AI by building intelligent systems that learn continuously from real-world execution. We transform complex enterprise workflows and business context into dynamic, interactive environments where AI agents can safely learn, adapt, and improve. By combining high-fidelity simulation environments with state-of-the-art training loops, we build specialized models that solve long-horizon, complex tasks with unmatched reliability. We're building the most talent-dense AI team in Europe to make this happen.


You'll develop post-training methods and own the performance of the systems that run them. The work spans learning algorithms, distributed training, and GPU performance, with responsibility for both model improvements and efficient experimentation.


What you'll work on


  • Implement and improve supervised fine-tuning, preference optimization, and RL methods such as PPO, GRPO and SDPO.


  • Own the training loop from rollout generation and reward computation through policy updates, weight synchronization, and checkpointing.


  • Improve throughput, GPU utilization, and memory efficiency through batching, parallelism, communication overlap, and asynchronous execution.


  • Profile bottlenecks and investigate numerical instability, stale-policy effects, and training-inference discrepancies. Verify that optimizations preserve learning behavior.


  • Occasionally deploy models in client environments and hill-climb alongside their teams: analyze real failures, refine rewards and training data, and iterate on task success, latency, and inference cost.



Relevant technologies


  • Training: PyTorch, Miles, NVIDIA Molt, verl, FSDP, and Megatron-LM


  • Rollout generation: vLLM, SGLang, Ray and asynchronous execution frameworks.


  • Performance/Kernels: NCCL, CUDA, Triton and or CuTE


  • Profiling: PyTorch Profiler, NVIDIA Nsight Systems



About you

You bring strong programming skills, quantitative reasoning, and an interest in how learning algorithms interact with the systems underneath them. You can move from an experiment to a profiler trace and investigate what the evidence shows.


We welcome both experienced specialists and generalists who learn exceptionally quickly. Experience with every method or framework above is not required.


What we offer


  • Competitive compensation and equity.


  • A fast-paced environment combining frontier research with impactful real-world applications.


  • Visa sponsorship and relocation support for candidates joining us in Paris or London.


  • A flexible hybrid setup, with a preference for working together in person.



Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Systems Engineer — RL & Post-Training Pipelines
ML Systems Engineer — RL & Post-Training Pipelines

Remanence • Paris (TX)

Hybrid
USD 135,000 - 203,000
Visa sponsorship
Relocation support
Hybrid setup
Member of Technical Staff, Post-Training
Member of Technical Staff, Post-Training

Goaly • Menlo Park (CA), Northern (KY)

On-site
USD 150,000 - 230,000
Meals and office benefits
Visa sponsorship
Member of Technical Staff — Environments / Evals
Member of Technical Staff — Environments / Evals

Remanence • Paris (TX)

Hybrid
USD 102,000 - 147,000
Visa sponsorship
Relocation support
Hybrid work
Member of Technical Staff, Post-Training & Applied Research
Member of Technical Staff, Post-Training & Applied Research

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Member of Technical Staff - ML Infrastructure Engineer, Post-training
Member of Technical Staff - ML Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 200,000 - 350,000
Health, vision, dental benefits
401K match
Lunch provided onsite
+2
Forward Deployed Engineer, Lead - LLM Post-training
Forward Deployed Engineer, Lead - LLM Post-training

reflectionai • New York (NY), California (MO)

On-site
USD 180,000 - 280,000
Top-tier compensation
Stock options
Health & wellness
+2
Member of Technical Staff — Infrastructure
Member of Technical Staff — Infrastructure

Remanence • Paris (TX)

Hybrid
USD 124,000 - 186,000
Visa sponsorship
Relocation support
Hybrid work setup
+1
Research Engineer / Research Scientist, RL Frontiers
Research Engineer / Research Scientist, RL Frontiers

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 500,000 - 850,000
Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research Engineer - Distributed Training
Research Engineer - Distributed Training

Prime Intellect • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 350,000