Lead Research Engineer / Research Scientist - Post-Training, Reinforcement Learning & Training Systems

Embodied AI

Lausanne

Hybrid

CHF 140.000 - 230.000

Vollzeit

Vor 11 Tagen
Bewerbungsgenerator

Eine komplette Bewerbung in einer Minute — maßgeschneiderter Lebenslauf und Anschreiben, versandbereit.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Embodied AI is seeking a Lead Research Engineer or Research Scientist to own the full post-training stack for production models. You will design, implement, and operate supervised fine-tuning, reinforcement learning, and policy distillation across multiple machines and accelerators.

You will guide researchers and engineers while staying hands-on. The role emphasizes scalable rollout systems, multilingual support, and end-to-end training reliability.

Qualifikationen

  • Proven experience in large-scale language-model training or post-training.
  • Experience with multi-machine, multi-accelerator workloads.
  • Proficiency in Python, PyTorch, autograd, mixed precision, and distributed execution.

Aufgaben

  • Own end-to-end post-training pipeline from pretrained checkpoints to production candidate.
  • Set technical direction for post-training and RL work.
  • Design and execute SFT (full-parameter and parameter-efficient).
  • Implement RLHF, RLAIF, reward models, verifiers, and related methods.
  • Develop scalable rollout systems and multi-stage curricula.

Kenntnisse

Python
PyTorch
Distributed training
Autodiff
Mixed precision
Reinforcement learning
RLHF
CUDA/NCCL
DeepSpeed
Megatron-Core

Tools

PyTorch Distributed
FSDP
DeepSpeed
Megatron-Core
Hugging Face Transformers
CUDA
Kubernetes
Ray

Jobbeschreibung

About the role

We are looking for a Lead Research Engineer or Research Scientist to own and lead the training and optimisation side of our complete post‑training stack. Starting from pretrained checkpoints, you will design, implement, scale, and operate the methods required to produce capable, reliable, and controllable production models. Your scope will include supervised fine‑tuning, preference optimisation, reinforcement learning, reward and verifier integration, policy distillation or consolidation, and distributed training. You will be expected to set technical direction in these areas, make key design decisions, and help guide the work of other researchers and engineers while remaining deeply hands‑on. This is not a single‑GPU fine‑tuning or adapter‑only role. You should be comfortable operating training workloads where memory, communication, rollout generation, hardware topology, and fault recovery must be designed together.

You will:
  • Own the end‑to‑end post‑training pipeline from pretrained checkpoint to production candidate.
  • Set technical direction and priorities for post‑training and reinforcement‑learning work.
  • Design and execute full‑parameter and parameter‑efficient SFT.
  • Implement preference optimisation, RLHF, RLAIF, reinforcement learning with verifiable rewards, and related methods.
  • Develop training strategies for reasoning, coding, tool use, multilingual behaviour, and long‑horizon agent tasks.
  • Integrate reward models, verifiers, critics, graders, and process‑or‑outcome‑based rewards.
  • Build scalable rollout‑generation systems for iterative and on‑policy training.
  • Design multi‑stage curricula combining SFT, reinforcement learning, rejection sampling, distillation, and policy consolidation.
  • Scale training across multiple machines and accelerators using appropriate combinations of data, tensor, pipeline, sequence, context, or expert parallelism.
  • Select sharding, precision, checkpointing, optimiser, batch‑size, sequence‑length, and activation‑recomputation strategies.
  • Estimate memory, communication, throughput, rollout capacity, and compute requirements before launching major runs.
  • Profile and improve accelerator utilisation, communication efficiency, data loading, and end‑to‑end training time.
  • Diagnose numerical instability, communication failures, out‑of‑memory errors, stragglers, checkpoint issues, and convergence regressions.
  • Investigate reward hacking, entropy collapse, KL drift, stale rollouts, mode collapse, grader exploitation, and benchmark overfitting.
  • Build reliable checkpointing, recovery, monitoring, and reproducibility procedures.
  • Collaborate closely with data, evaluation, infrastructure, and inference teams.
  • Provide technical guidance and mentorship to other researchers and engineers working on training and post‑training.
  • Contribute clean, tested code, technical reports, and operational runbooks.
We are looking for demonstrated experience in most of the following areas:
  • Ownership of large‑scale language‑model training or post‑training runs across multiple machines and accelerators.
  • Experience with workloads for which straightforward single‑node training or pure data parallelism was insufficient.
  • Deep proficiency with Python, PyTorch, autograd, mixed precision, optimisation, and distributed execution.
  • Practical experience with PyTorch Distributed, FSDP, DeepSpeed, Megatron‑Core, or an equivalent framework.
  • Ability to select parallelism and sharding strategies based on model, sequence, memory, and network constraints.
  • Strong understanding of SFT, preference optimisation, reinforcement learning, reward modelling, KL regularisation, sampling, and training stability.
  • Experience operating high‑throughput inference or rollout systems as part of a training loop.
  • Ability to debug across model code, distributed communication, numerical optimisation, data, and infrastructure.
  • Strong experimental design and the ability to distinguish algorithmic improvements from evaluation or systems artefacts.
  • Experience building reliable, observable, and reproducible research software.
  • Personal ownership of consequential decisions affecting a substantial training programme.
  • Demonstrated ability to provide technical direction, review complex design choices, and raise the technical level of a team.
Relevant stack:
  • Python and PyTorch.
  • PyTorch Distributed and FSDP.
  • DeepSpeed, Megatron‑Core, or comparable frameworks.
  • Hugging Face Transformers.
  • CUDA and NCCL.
  • vLLM, SGLang, or similar rollout engines.
  • Ray, Slurm, Kubernetes, or comparable orchestration systems.
  • MLflow or Weights & Biases.
  • Docker, GCP, GitLab CI, profiling, monitoring, and pytest.
  • Experience with CUDA or Triton, long‑context training, sparse models, asynchronous RL, stateful agent environments, distillation, or deployment‑aware post‑training would be especially valuable.
You may be a strong fit if you:
  • Enjoy working at the intersection of model research and distributed systems.
  • Can move from paper reproduction to reliable scaled implementation.
  • Are comfortable taking responsibility for expensive and operationally demanding experiments.
  • Approach failures methodically across algorithms, data, numerical stability, and infrastructure.
  • Care about held‑out capability and reliability, not only training loss or reward.
  • Are comfortable setting technical direction while remaining deeply hands‑on.
  • Want meaningful ownership of a complete model programme.
Location and work style:
  • We offer full‑time employment in Switzerland.
  • Remote work is supported.
  • The team gathers approximately twice per month in a Swiss office.
  • Exceptional candidates elsewhere in Europe may be considered.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Research Engineer / Research Scientist - Post-Training, Reinforcement Learning & Training Systems
Senior Research Engineer / Research Scientist - Post-Training, Reinforcement Learning & Training Systems

Embodied AI • Lausanne

Hybrid
CHF 150.000 - 210.000
AI Research Scientist
AI Research Scientist

Embodied AI • Lausanne

Hybrid
CHF 120.000 - 170.000
AI Research Scientist
AI Research Scientist

Giotto.ai • Lausanne

Hybrid
CHF 140.000 - 210.000
Apertus Engineer: Post-training
Apertus Engineer: Post-training

ETH Zürich • Zürich

Hybrid
CHF 110.000 - 170.000
Access to HPC infrastructure
Open-source collaboration
Apertus Engineer: Post-training 100%
Apertus Engineer: Post-training 100%

ETH Zürich • Zürich

Hybrid
CHF 110.000 - 150.000
Remote work options
Professional development
Open‑source projects
Reward Model Research Engineer
Reward Model Research Engineer

Rapidata AG • Zürich

Vor Ort
CHF 120.000 - 180.000
Competitive salary & equity
Mountain-view office near Sihlcity
Hardware budget
+2
Apertus Engineer: Post-training
Apertus Engineer: Post-training

Master in Integrated Building Systems ETH Zürich • Zürich

Hybrid
CHF 120.000 - 170.000
Flexible working arrangements
Professional development opportunities
Access to cutting-edge HPC
Staff / Principal Machine Learning Engineer, Serving - Switzerland
Staff / Principal Machine Learning Engineer, Serving - Switzerland

Inworld • Schweiz

Remote
CHF 90.000 - 120.000
Member of Technical Staff - AI Research Engineer
Member of Technical Staff - AI Research Engineer

Sarah Smith • Zürich

Hybrid
CHF 90.000 - 120.000
Visa sponsorship
Flexible PTO
Regular offsites and in-person events
Founding Member of Technical Staff Zurich ML Systems
Founding Member of Technical Staff Zurich ML Systems

Fainite AG • Zürich

Hybrid
CHF 180.000 - 280.000