Reward Model Research Engineer

Rapidata AG

Zürich

Vor Ort

CHF 120.000 - 180.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Benefits dieser Stelle

Competitive salary & equity
Mountain-view office near Sihlcity
Hardware budget
Unlimited snacks and drinks
Growth opportunities

Zusammenfassung

Rapidata AG in Zürich is seeking a Reward Model Research Engineer to design, train, and productionize reward models that transform large-scale human feedback into reliable signals for RLHF and DPO. You will tackle PRMs, generalization across models, and inference-time guidance in a live, global-scale system.

You will collaborate with the data platform team to reduce bottlenecks in training and validation, while translating research into production-ready pipelines and benchmarking for robustness

Qualifikationen

  • Hands-on research experience building reward models for LLMs or diffusion models.
  • Solid understanding of reinforcement learning fundamentals and RLHF/DPO training pipelines.
  • Practical experience with inference-time guidance techniques and evaluating multi-step reasoning.
  • Strong Python and PyTorch skills.
  • Experience turning research prototypes into production-ready systems.
  • Solid statistical and mathematical foundation.
  • Excellent English communication skills, both written and oral.

Aufgaben

  • Design, train, and evaluate reward models, including process reward models (PRMs), from large-scale human preference data.
  • Build and improve inference-time guidance methods to make reward-guided generation more robust.
  • Develop training setups that improve generalization across diverse generator/policy models.
  • Translate reward modeling research into production-ready pipelines that plug directly into Rapidata's RLHF/DPO data flywheel.
  • Collaborate with the data platform team to design data collection and active learning strategies that reduce reward model training bottlenecks.
  • Rigorously benchmark reward models for accuracy, robustness, and reliability before deployment.
  • Communicate findings clearly to both technical and non-technical stakeholders, including partner AI labs.

Kenntnisse

Reward modeling
Reinforcement learning
Python
Communication

Ausbildung

MSc/PhD in a relevant field

Tools

PyTorch

Jobbeschreibung

About Rapidata

Rapidata provides an API to humans that is revolutionizing the data generation and annotation industry. We deliver highly scalable, extremely fast human feedback that fuels the AI systems of the future, powering RLHF and DPO training data collection at internet speed for frontier AI labs. Our network reaches over 20 million active annotators across 192 countries, distributing micro-tasks ("Rapids") and returning verified labels in near real-time.

The Role

We\'re looking for a Reward Model Research Engineer to design, train, and productionize the reward models work alongside Rapidata\'s real-time human feedback for high-quality training signal for RLHF and DPO. This is a research-to-production role: you\'ll work on the modeling problems that determine whether noisy, large-scale human preference data becomes a reliable reward signal, including process reward models (PRMs) that score multi-step reasoning.

You\'ll be closing the loop between our human-in-the-loop data collection infrastructure and the reward models that consume it and accompany it, tackling generalization across generator models, robustness to annotator noise, and inference-time guidance techniques, then validating your models in a live system operating at global scale.

What You\'ll Do
  • Design, train, and evaluate reward models, including process reward models (PRMs), from large-scale human preference data

  • Build and improve inference-time guidance methods to make reward-guided generation more robust

  • Develop training setups that improve generalization across diverse generator/policy models

  • Translate reward modeling research into production-ready pipelines that plug directly into Rapidata\'s RLHF/DPO data flywheel

  • Collaborate with the data platform team to design data collection and active learning strategies that reduce reward model training bottlenecks

  • Rigorously benchmark reward models for accuracy, robustness, and reliability before deployment

  • Communicate findings clearly to both technical and non-technical stakeholders, including partner AI labs

What We\'re Looking For
  • Hands-on research experience building reward models for LLMs or diffusion models, ideally through an MSc/PhD thesis or equivalent applied research

  • Solid understanding of reinforcement learning fundamentals and RLHF/DPO training pipelines

  • Practical experience with inference-time guidance techniques and evaluating multi-step reasoning

  • Strong Python and deep learning framework skills (PyTorch)

  • Experience turning research prototypes into validated, production-ready systems

  • Solid statistical and mathematical foundation

  • Excellent English communication skills, both oral and written

Nice-to-Have
  • Experience with agentic system safety, guardrails, or LLM-based tool-calling agents

  • Publications or open-source contributions in reward modeling, reasoning, or reinforcement learning

  • Experience with hierarchical or model-based RL

  • Based in or willing to relocate to Zürich

What We Offer
  • Competitive salary and equity in a startup with strong growth, IP, and backing from top-tier VCs

  • Opportunity to join a fast-growing startup early, with an outsized opportunity to shape where the company goes

  • Opportunities for personal and professional growth as our team expands

  • Fun and open (startup) culture

  • Spacious mountain-view office near Sihlcity, Zürich, with terrace, table tennis, pizza oven, hammock, and BBQ

  • Hardware budget tailored to your preferences

  • Unlimited snacks and drinks of your choice

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Reward Model Research Engineer: From Research to Production
Reward Model Research Engineer: From Research to Production

Rapidata AG • Zürich

Vor Ort
CHF 120.000 - 180.000
Competitive salary & equity
Mountain-view office near Sihlcity
Hardware budget
+2
Backend Software Engineer
Backend Software Engineer

Next Matter • Zürich

Vor Ort
CHF 90.000 - 120.000
Competitive salary and equity
Unlimited snacks and drinks
Large terrace with recreational facilities
Senior Systems Engineer / Ad Tech Engineer Zürich, Switzerland
Senior Systems Engineer / Ad Tech Engineer Zürich, Switzerland

Rapidata AG • Zürich

Vor Ort
CHF 90.000 - 130.000
Competitive salary and equity
Personal and professional growth opportunities
Fun and open startup culture
+2
AI Platform Engineer
AI Platform Engineer

Rapidata AG • Zürich

Vor Ort
CHF 80.000 - 120.000
Competitive salary and equity
Opportunity for personal growth
Spacious office with mountain views
+1
Senior Game/Product Designer (behavioral)
Senior Game/Product Designer (behavioral)

Next Matter • Zürich

Vor Ort
CHF 90.000 - 130.000
Spacious Zürich Center office with MTB
Hardware budget tailored to your needs
Unlimited snacks and drinks
AI Research Scientist
AI Research Scientist

Giotto.ai • Lausanne

Hybrid
CHF 140.000 - 210.000
AI Platform Engineer
AI Platform Engineer

Next Matter • Zürich

Vor Ort
CHF 80.000 - 100.000
Competitive salary and equity
Personal and professional growth opportunities
Fun startup culture
+3
Senior Systems Engineer / Ad Tech Engineer
Senior Systems Engineer / Ad Tech Engineer

Next Matter • Zürich

Vor Ort
CHF 90.000 - 120.000
Competitive salary and equity
Hardware budget tailored to preferences
Unlimited snacks and drinks
Backend Software Engineer Zürich, Switzerland
Backend Software Engineer Zürich, Switzerland

Rapidata AG • Zürich

Vor Ort
CHF 120.000 - 200.000
Competitive salary plus equity
Unlimited snacks and beverages
Office with mountain views
UX/UI Product Designer (Behavioral Design) Zürich, Switzerland
UX/UI Product Designer (Behavioral Design) Zürich, Switzerland

Rapidata AG • Zürich

Vor Ort
CHF 80.000 - 120.000
Competitive salary and equity
Opportunities for personal and professional growth
Spacious office with mountain views
+1