Member of Technical Staff, Reinforcement Learning

Inception

San Francisco (CA)

On-site

USD 180,000 - 260,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Inception seeks experienced scientists and engineers with deep expertise in post-training large language models through reinforcement learning. You will design and implement RL training pipelines for diffusion LLMs, develop reward modeling strategies, and build the algorithms that align model behavior with human intent at scale.

The role focuses on designing RL training pipelines, building reward models, and advancing alignment techniques for enterprise-grade LLMs, with emphasis on stability,

Qualifications

  • PhD or equivalent research/engineering experience in ML.
  • Proficient with PyTorch and distributed training.
  • Strong knowledge of transformers, LLMs, and RLHF workflows.

Responsibilities

  • Design, develop, and optimize RL training pipelines (PPO, DPO, RLHF) for diffusion-based LLMs.
  • Build and iterate reward models and evaluation of reward quality.
  • Implement approaches for fine-tuning and scaling generative AI models.
  • Work on data preprocessing, model evaluation, and enterprise alignment.
  • Research techniques for controlled text generation and constraint satisfaction.
  • Improve training stability, efficiency, and reproducibility of RL workloads.

Skills

RLHF experience
PPO/DPO familiarity
Transformer models
Distributed training
PyTorch experience

Education

BS/MS/PhD in CS

Tools

vLLM
TensorRT
SGLang

Job description

The Role

We seek experienced scientists and engineers with deep expertise in post-training large language models through reinforcement learning. You will design and implement RL training pipelines for our diffusion LLMs, develop reward modeling strategies, and build the algorithms that align model behavior with human intent at scale.

Key Responsibilities
  • Design, develop, and optimize RL training pipelines (PPO, DPO, RLHF, and novel approaches) for diffusion-based LLMs.
  • Build and iterate on reward models, reward shaping strategies, and evaluation of reward quality.
  • Implement innovative approaches for fine-tuning and scaling generative AI models.
  • Work on data preprocessing pipelines, model evaluation, and alignment to enterprise use cases.
  • Research and implement techniques for controlled text generation and constraint satisfaction.
  • Improve training stability, efficiency, and reproducibility of RL workloads.
Qualifications
  • BS/MS/PhD in Computer Science or a related field (or equivalent experience).
  • At least 2 years of experience working on ML projects in PyTorch (or equivalent), preferably in a research lab or engineering role.
  • Excellent familiarity with transformers and core LLM concepts (autoregressive pretraining, instruction tuning, in-context learning, KV caching).
  • Hands-on experience with reinforcement learning from human feedback (RLHF), PPO, DPO, or related post-training methods.
  • Familiarity with training and inference in diffusion models.
  • Experience training deep learning models at scale in distributed computing environments.
Preferred Skills
  • Extensive experience training transformer-based language models from scratch.
  • Experience designing and implementing reward models or preference learning systems.
  • Knowledge of advanced training techniques (mixed precision, gradient accumulation, etc.).
  • Background in optimization theory and neural network architecture design.
  • Experience with LLM serving frameworks like vLLM, SGLang, or TensorRT.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Training
Member of Technical Staff, Training

Inception • San Francisco (CA)

On-site
USD 180,000 - 230,000
Member of Technical Staff, RL Infra
Member of Technical Staff, RL Infra

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000
Staff Scientist: RL & Reward Modeling for Diffusion LLMs
Staff Scientist: RL & Reward Modeling for Diffusion LLMs

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
Research Scientist, LLM Evaluation & Post-Training
Research Scientist, LLM Evaluation & Post-Training

OneForma • United States

Hybrid
USD 140,000 - 210,000
RL Research Scientist — Large Language Models
RL Research Scientist — Large Language Models

Brahma Consulting Group • Seattle (WA)

On-site
USD 150,000 - 230,000
Machine Learning Engineer, LLM Post-Training
Machine Learning Engineer, LLM Post-Training

GoTo Meeting • Mountain View (CA)

On-site
USD 150,000 - 230,000
Health, dental, and vision care for you and your family
Top-tier 401(K) plan with company matching
Paid time off and paid holidays
+2
Research Scientist - Post-training / RL
Research Scientist - Post-training / RL

Epsilon Health • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Researcher
Machine Learning Researcher

Brahma Consulting Group • San Francisco (CA)

On-site
USD 150,000 - 230,000
Member of Technical Staff, Post-training
Member of Technical Staff, Post-training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Research Scientist / Engineer – Reinforcement Learning Infrastructure

Luma • Redwood City (CA)

On-site
USD 200,000 - 300,000