Post-Training Research Scientist

Two Sigma

New York (NY)

On-site

USD 180,000 - 240,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Two Sigma seeks a researcher-focused engineer to lead post-training efforts for language models applied to financial time series and quantitative reasoning. You will design RLHF, DPO, and related methods at scale, managing substantial compute budgets and building data/reward infrastructures on financial datasets.

You will collaborate with quant researchers to align post-training work with finance applications, driving a strong research agenda and ensuring smooth production integration.

Qualifications

  • Shipped post-training systems in production (RLHF, DPO, RLAIF) with alignment goals.
  • Strong understanding of post-training methods and alignment research.
  • Hands-on experience with PyTorch/JAX and distributed frameworks.

Responsibilities

  • Lead post-training efforts for LLMs applied to financial time series and quantitative reasoning.
  • Design and execute RLHF, DPO, and related alignment methods at scale, including deployment of substantial compute budgets.
  • Build infrastructure for preference data collection, reward modeling, and policy optimization on financial datasets
  • Drive research agenda connecting post-training methods to quantitative finance applications
  • Collaborate with quant researchers to define task distributions and evaluation frameworks
  • Unblock production systems dependent on post-training capabilities

Skills

PyTorch/JAX
Distributed training
Experiment design
Budgeting & compute management

Education

BS or MS in STEM

Tools

DeepSpeed
FSDP
Multi-node GPU clusters

Job description

Responsibilities
  • Lead post-training efforts for LLMs applied to financial time series and quantitative reasoning
  • Design and execute RLHF, DPO, and related alignment methods at scale, including deployment of substantial compute budgets (O($100mm))
  • Build infrastructure for preference data collection, reward modeling, and policy optimization on financial datasets
  • Drive research agenda connecting post-training methods to quantitative finance applications
  • Collaborate with quant researchers to define task distributions and evaluation frameworks
  • Unblock production systems dependent on post-training capabilities
Qualifications
  • BS or equivalent work experience in Science, Technology, Engineering or Math (an MS is a plus).
  • Minimum 1 year of experience required; 1-10 years of experience preferred (ideally 1-5 years) at a frontier AI lab (OpenAI, Anthropic, DeepMind, Meta FAIR, or equivalent)
  • Shipped post-training systems in production: RLHF, DPO, RLAIF, or related methods
  • Deep understanding of distributed training infrastructure: multi-node GPU clusters, training stability, checkpointing
  • Track record managing large-scale compute: budgeting, experiment design, ablations
  • Publications or demonstrated expertise in alignment, preference learning, or reward modeling
  • Hands-on implementation skills: PyTorch/JAX, distributed frameworks (DeepSpeed, FSDP, etc.)
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

RESEARCHER, POST-TRAINING
RESEARCHER, POST-TRAINING

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Lead Post-Training Research Scientist, Financial AI
Lead Post-Training Research Scientist, Financial AI

Two Sigma • New York (NY)

On-site
USD 180,000 - 240,000
Member of Technical Staff - Research Engineer, Post-training
Member of Technical Staff - Research Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive cash and equity (>90th pct
Ownership and autonomy
Lunch onsite
+4
Research Scientist – RL Post-Training for Agents
Research Scientist – RL Post-Training for Agents

Rnb Consultancy • San Francisco (CA)

On-site
USD 250,000 - 500,000
Visa sponsorship
Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)
Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)

Hippocratic-Ai • Menlo Park (CA)

On-site
USD 230,000 - 290,000
Machine Learning Engineer, LLM Post-Training
Machine Learning Engineer, LLM Post-Training

GoTo Meeting • Mountain View (CA)

On-site
USD 150,000 - 230,000
Health, dental, and vision care for you and your family
Top-tier 401(K) plan with company matching
Paid time off and paid holidays
+2
Research Scientist - Post-training / RL
Research Scientist - Post-training / RL

Epsilon • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff, Reinforcement Learning
Member of Technical Staff, Reinforcement Learning

Inception • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior Staff Research Scientist/Engineer, Agentic Post Training, DeepMind
Senior Staff Research Scientist/Engineer, Agentic Post Training, DeepMind

AI Chopping Block, Inc. • Mountain View (CA), Northern (KY)

Hybrid
USD 262,000 - 364,000
Equity
Bonus target
Benefits package
Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)
Applied Scientist, Reinforcement Learning (Mid, Senior, Staff)

Hippocratic AI • Menlo Park (CA)

On-site
USD 180,000 - 240,000