Research Scientist - Post-training / RL

Epsilon Labs, Inc.

San Francisco (CA)

On-site

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Epsilon Labs, Inc. is seeking a Research Scientist with deep expertise in post-training and reinforcement learning to advance multimodal models for clinical radiology use.

You will own stages after pretraining, including supervised fine-tuning, reward modeling, and inference-time strategies, with a focus on real-world patient outcomes. The role involves designing verification-friendly rewards, RLHF pipelines, and scalable ML systems on large medical imaging datasets, contributing to high-impact

Qualifications

  • 6+ years of academia/industry experience in reinforcement learning, post-training, or multimodal machine learning
  • Deep expertise in post-training large language or vision-language models
  • Strong foundation in modern post-training and reinforcement learning techniques including multi-reward objectives
  • Proficiency in PyTorch or JAX, with multi-GPU/distributed training experience
  • Experience with reinforcement learning infrastructure at scale and related frameworks
  • Experience with autoregressive modeling and instruction tuning
  • Strong software engineering skills and production-quality code

Responsibilities

  • Design reinforcement learning with verifiable rewards for report generation and grounding
  • Extend RL to noisy objectives using learned reward models and radiologist feedback
  • Run GRPO-family algorithms with multi-reward objectives and diagnose issues
  • Train explicit reward models conditioned on image data with supervision
  • Train chain‑of‑thought reasoning over image regions with grounding
  • Develop multimodal tool use and propagate credit across trajectories
  • Create inference-time methods like best‑of‑N sampling and verifier-guided decoding
  • Tune output styling to institutional reporting conventions
  • Stay current with cutting-edge research and publish results
  • Lead research through conferences and technical blogs

Skills

Reinforcement learning
Post-training
Multimodal ML
Grounded report generation
RLHF
Reward modeling
Inference-time strategies
Tool use in ML
Reasoning over images
Model optimization

Tools

PyTorch
JAX
vLLM
OpenRLHF
TRL

Job description

About Us

We're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting‑edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product‑market fit with a substantial customer pipeline already in place.

Role Overview

We're seeking a Research Scientist with deep expertise in post‑training and reinforcement learning to join our ML Research team. You'll be at the forefront of developing and deploying state‑of‑the‑art multimodal models for clinical use in radiology settings. This role owns every stage after pretraining: supervised fine‑tuning, reward modeling, reinforcement learning against verifiable and learned reward signals, reasoning and tool‑use training, and inference‑time strategy. You'll work with one of the largest and most diverse medical imaging datasets in the industry, advancing the state‑of‑the‑art in grounded report generation, reward design, and inference‑time reasoning while maintaining the clinical rigor required for healthcare deployment.

Key Responsibilities
  • Design reinforcement learning with verifiable rewards for report generation, including clinical label and entity‑relation matching, grounding IoU, measurement accuracy, and reporting schema compliance.

  • Extend reinforcement learning to unverifiable and noisy objectives such as report quality and clinical usefulness, using learned reward models and radiologist feedback pipelines (RLHF) built on expert preferences and report edits.

  • Run GRPO‑family algorithms with complex multi‑reward objectives, tuning reward composition and diagnosing reward hacking, entropy collapse, and diversity loss.

  • Train explicit reward models, including multimodal reward models conditioned on the image, with both outcome and process supervision.

  • Train chain‑of‑thought reasoning over image regions, including evidence localization and verification loops that keep reasoning grounded in the image rather than in language priors.

  • Train multimodal tool use — windowing, zoom and crop, detector and segmentation calls, prior study retrieval — with credit assignment across multi‑turn trajectories.

  • Develop inference‑time methods including best‑of‑N sampling against reward models and grounding‑aware decoding, and distill the resulting gains back into the policy.

  • Tune output stylization to institutional reporting conventions, keeping style rewards separated from clinical content rewards.

  • Stay current with cutting‑edge research in reinforcement learning, reward modeling, and multimodal post‑training.

  • Drive research and technical excellence through conference publications and technical blog posts, establishing best practices for post‑training medical VLMs at scale.

Qualifications
  • 6+ years of academia/industry experience in reinforcement learning, post‑training, or multimodal machine learning

  • Deep expertise in post‑training large language or vision‑language models (e.g., Qwen‑VL, InternVL, LLaVA, or similar architectures)

  • Strong foundation in modern post‑training and reinforcement learning techniques including:

    • Group‑relative policy optimization and its successors (GRPO, DAPO, GSPO, CISPO) with multi‑reward objectives

    • Reinforcement learning with verifiable rewards, and with noisy, sparse, or learned reward signals

    • Reward model training: pairwise and generative reward models, outcome and process supervision

    • Preference optimization methods (DPO, IPO, ORPO, KTO) and RLHF

    • Inference‑time compute scaling, including best‑of‑N sampling and verifier‑guided decoding

  • Practical experience diagnosing and mitigating reward hacking and reward over‑optimization

  • Track record of implementing complex models from research papers and adapting them to new domains

  • Proficiency in PyTorch or JAX, with experience training large models on multi‑GPU/distributed systems

  • Experience with reinforcement learning infrastructure at scale, including rollout generation (vLLM, SGLang) and frameworks such as verl, TRL, or OpenRLHF

  • Experience with autoregressive language modeling and instruction tuning

  • Strong software engineering skills and ability to write production‑quality code

Preferred Qualifications
  • Publications at top‑tier conferences (NeurIPS, ICML, ICLR, CVPR, ACL, EMNLP, MICCAI)

  • Hands‑on experience with medical imaging applications, particularly radiology report generation

  • Experience with agentic or multi‑turn reinforcement learning, including credit assignment over tool‑use trajectories

  • Experience with grounded generation tasks (visual grounding, referring expression comprehension)

  • Knowledge of evaluation methodologies for long‑form generation, including factuality assessment and hallucination detection

  • Experience mitigating catastrophic forgetting of supervised capabilities during reinforcement learning

  • Familiarity with clinical NLP and medical knowledge representation

  • Experience with model interpretability, explainability, and uncertainty quantification in safety‑critical applications

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Scientist - VLM Pretraining
Research Scientist - VLM Pretraining

Epsilon Labs, Inc. • San Francisco (CA)

On-site
USD 180,000 - 260,000
Research Scientist - VLM Pretraining
Research Scientist - VLM Pretraining

Epsilon Health • San Francisco (CA)

On-site
USD 180,000 - 250,000
Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)
Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)

Hippocratic-Ai • Menlo Park (CA)

On-site
USD 230,000 - 290,000
Research Engineer - Data Quality & Evals
Research Engineer - Data Quality & Evals

Epsilon Health • San Francisco (CA)

On-site
USD 120,000 - 170,000
Research Scientist, Medical Imaging & RL Post-Training
Research Scientist, Medical Imaging & RL Post-Training

Epsilon Labs, Inc. • San Francisco (CA)

On-site
USD 180,000 - 240,000
Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)
Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)

Hippocratic AI • Menlo Park (CA)

On-site
USD 180,000 - 240,000
RESEARCHER, POST-TRAINING
RESEARCHER, POST-TRAINING

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Research Scientist - Vision Foundation Models
Research Scientist - Vision Foundation Models

Epsilon Labs, Inc. • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Competitive cash and equity compensation (>90th percentile)
Ownership and autonomy
Health, vision, dental benefits
+4
Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2