Lead Research Scientist - Post-Training

Giotto.ai

Lausanne

Hybrid

CHF 150,000 - 210,000

Full time

6 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Giotto.ai in Switzerland seeks a Lead Research Scientist to own and lead the post-training stack, including supervised fine-tuning, RLHF, and verification components. You will design scalable training systems, manage memory and parallelism, and mentor researchers while remaining hands-on.

The role emphasizes end-to-end pipelines, multi-machine training, and collaboration with data, evaluation, and infrastructure teams to deliver reliable, controllable production models.

Qualifications

  • Experience leading large-scale language-model training or post-training runs across multiple machines.
  • Proven ability to design and run end-to-end post-training pipelines in production.
  • Strong Python and PyTorch proficiency with distributed execution expertise.
  • Hands-on experience with DeepSpeed, Megatron-Core, or equivalent frameworks.

Responsibilities

  • Own the end-to-end post-training pipeline from pretrained checkpoint to production candidate.
  • Set technical direction for post-training and reinforcement-learning work.
  • Design and execute full-parameter and parameter-efficient SFT.
  • Implement preference optimisation, RLHF, and related methods.
  • Develop training strategies for reasoning, tool use, multilingual behaviour, and long-horizon tasks.
  • Integrate reward models and verifiers; build scalable rollout-generation systems.
  • Scale training across multiple machines and accelerators with suitable parallelism.

Skills

LLM training
Python
PyTorch
Distributed training
DeepSpeed
Transformers
CUDA/NCCL
Reinforcement learning
Experiment design
Parallelism strategies

Tools

DeepSpeed
Megatron-Core
Hugging Face Transformers
CUDA
NCCL
Docker
GCP
Weights & Biases

Job description

Giotto.ai is a Switzerland-based AI company building intelligence systems for Switzerland and Europe.

Our mission is to enable governments and enterprises to retain control over the AI systems they use without compromising access to advanced reasoning capabilities. Giotto combines portable, configurable models with an AI operating system, integrating open and proprietary weights, datasets, tools, and deployment components.

About the role

We are looking for a Lead Research Scientist to own and lead the training and optimisation side of our complete post-training stack.

Starting from pretrained checkpoints, you will design, implement, scale, and operate the methods required to produce capable, reliable, and controllable production models.

Your scope will include supervised fine-tuning, preference optimisation, reinforcement learning, reward and verifier integration, policy distillation or consolidation, and distributed training.

You will be expected to set technical direction in these areas, make key design decisions, and help guide the work of other researchers and engineers while remaining deeply hands‑on.

This is not a single-GPU fine-tuning or adapter-only role. You should be comfortable operating training workloads where memory, communication, rollout generation, hardware topology, and fault recovery must be designed together.

You will:
  • Own the end-to-end post-training pipeline from pretrained checkpoint to production candidate.
  • Set technical direction and priorities for post-training and reinforcement-learning work.
  • Design and execute full-parameter and parameter-efficient SFT.
  • Implement preference optimisation, RLHF, RLAIF, reinforcement learning with verifiable rewards, and related methods.
  • Develop training strategies for reasoning, coding, tool use, multilingual behaviour, and long-horizon agent tasks.
  • Integrate reward models, verifiers, critics, graders, and process- or outcome-based rewards.
  • Build scalable rollout-generation systems for iterative and on-policy training.
  • Design multi-stage curricula combining SFT, reinforcement learning, rejection sampling, distillation, and policy consolidation.
  • Scale training across multiple machines and accelerators using appropriate combinations of data, tensor, pipeline, sequence, context, or expert parallelism.
  • Select sharding, precision, checkpointing, optimiser, batch‑size, sequence‑length, and activation‑recomputation strategies.
  • Estimate memory, communication, throughput, rollout capacity, and compute requirements before launching major runs.
  • Profile and improve accelerator utilisation, communication efficiency, data loading, and end‑to‑end training time.
  • Diagnose numerical instability, communication failures, out‑of‑memory errors, stragglers, checkpoint issues, and convergence regressions.
  • Investigate reward hacking, entropy collapse, KL drift, stale rollouts, mode collapse, grader exploitation, and benchmark overfitting.
  • Build reliable checkpointing, recovery, monitoring, and reproducibility procedures.
  • Collaborate closely with data, evaluation, infrastructure, and inference teams.
  • Provide technical guidance and mentorship to other researchers and engineers working on training and post‑training.
  • Contribute clean, tested code, technical reports, and operational runbooks.
We are looking for demonstrated experience in most of the following areas:
  • Ownership of large‑scale language‑model training or post‑training runs across multiple machines and accelerators.
  • Experience with workloads for which straightforward single‑node training or pure data parallelism was insufficient.
  • Deep proficiency with Python, PyTorch, autograd, mixed precision, optimisation, and distributed execution.
  • Practical experience with PyTorch Distributed, FSDP, DeepSpeed, Megatron‑Core, or an equivalent framework.
  • Ability to select parallelism and sharding strategies based on model, sequence, memory, and network constraints.
  • Strong understanding of SFT, preference optimisation, reinforcement learning, reward modelling, KL regularisation, sampling, and training stability.
  • Experience operating high‑throughput inference or rollout systems as part of a training loop.
  • Ability to debug across model code, distributed communication, numerical optimisation, data, and infrastructure.
  • Strong experimental design and the ability to distinguish algorithmic improvements from evaluation or systems artefacts.
  • Experience building reliable, observable, and reproducible research software.
  • Personal ownership of consequential decisions affecting a substantial training programme.
  • Demonstrated ability to provide technical direction, review complex design choices, and raise the technical level of a team.

A PhD is not required. We value exceptional technical work, strong judgement, and demonstrated ownership.

  • Python and PyTorch.
  • PyTorch Distributed and FSDP.
  • DeepSpeed, Megatron‑Core, or comparable frameworks.
  • Hugging Face Transformers.
  • CUDA and NCCL.
  • vLLM, SGLang, or similar rollout engines.
  • MLflow or Weights & Biases.
  • Docker, GCP, GitLab CI, profiling, monitoring, and pytest.

Experience with CUDA or Triton, long-context training, sparse models, asynchronous RL, stateful agent environments, distillation, or deployment‑aware post‑training would be especially valuable.

You may be a strong fit if you:
  • Enjoy working at the intersection of model research and distributed systems.
  • Can move from paper reproduction to reliable scaled implementation.
  • Are comfortable taking responsibility for expensive and operationally demanding experiments.
  • Approach failures methodically across algorithms, data, numerical stability, and infrastructure.
  • Care about held‑out capability and reliability, not only training loss or reward.
  • Are comfortable setting technical direction while remaining deeply hands‑on.
  • Want meaningful ownership of a complete model programme.
Location and work style:

We offer full-time employment in Switzerland.

  • Remote work is supported.
  • The team gathers approximately twice per month in a Swiss office.
  • Exceptional candidates elsewhere in Europe may be considered.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Research Engineer / Research Scientist - Post-Training, Reinforcement Learning & Trainin[...]
Senior Research Engineer / Research Scientist - Post-Training, Reinforcement Learning & Trainin[...]

Giotto.ai • Lausanne

Remote
CHF 180,000 - 240,000
AI Research Scientist
AI Research Scientist

Embodied AI • Lausanne

On-site
CHF 120,000 - 170,000
LLM Research Scientist
LLM Research Scientist

Giotto.ai • Lausanne

Hybrid
CHF 120,000 - 190,000
Lead Research Engineer - Post-Training ML Systems (Remote)
Lead Research Engineer - Post-Training ML Systems (Remote)

Giotto.ai • Lausanne

On-site
CHF 180,000 - 240,000
Junior AI Researcher
Junior AI Researcher

Startupvalleys • Lausanne

Hybrid
CHF 65,000 - 110,000
Senior AI Research Engineer — Post-Training & Scale
Senior AI Research Engineer — Post-Training & Scale

Giotto.ai • Lausanne

Remote
CHF 180,000 - 240,000
Senior Post-Training AI Scientist (RL & Scaling)
Senior Post-Training AI Scientist (RL & Scaling)

Giotto.ai • Lausanne

Hybrid
CHF 150,000 - 210,000
Research Engineer
Research Engineer

Clera • Zürich

On-site
CHF 90,000 - 130,000
Equity
Visa sponsorship
Junior Machine Learning Engineer
Junior Machine Learning Engineer

Startupvalleys • Lausanne

Hybrid
CHF 90,000 - 120,000
Hybrid work model
Member of Technical Staff - AI Research Engineer
Member of Technical Staff - AI Research Engineer

Sarah Smith • Zürich

On-site
CHF 90,000 - 120,000
Visa sponsorship
Flexible PTO
Regular offsites and in-person events