Post-Training AI Research Scientist: RLHF for Finance

Two Sigma

New York (NY)

Hybrid

USD 165,000 - 300,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health insurance
401k match
Life & disability insurance
Gym access
Tuition reimbursement
Generous vacation & sick days
Hybrid work policy
Home office budget

Job summary

Two Sigma is advancing post-training with RLHF, DPO, and reward modeling to align LLMs with complex, multi-step financial workflows. You will define the research agenda, build scalable infrastructure, and guide evaluation frameworks for transforming quant research into production-grade AI capabilities.

The role combines training, fine-tuning, context management, and model evaluation, shaping both the post-training capability and the broader research direction of the team.

Qualifications

  • BS or equivalent in Science, Technology, Engineering or Math (an MS is a plus).
  • Minimum 1 year of experience required; 1-10 years preferred at a frontier AI lab (OpenAI, Anthropic, DeepMind, Meta FAIR, or equivalent).
  • Shipped post-training systems in production: RLHF, DPO, RLAIF, or related methods.
  • Deep understanding of distributed training infrastructure: multi-node GPU clusters, training stability, checkpointing.
  • Track record managing large-scale compute: budgeting, experiment design, ablations.
  • Publications or demonstrated expertise in alignment, preference learning, or reward modeling.
  • Hands-on implementation skills: PyTorch/JAX, distributed frameworks (DeepSpeed, FSDP, etc.)

Responsibilities

  • Lead post-training efforts for LLMs applied to financial time series and quantitative reasoning
  • Design and execute RLHF, DPO, and related alignment methods at scale, including deployment of substantial compute budgets (O($100mm))
  • Build infrastructure for preference data collection, reward modeling, and policy optimization on financial datasets
  • Drive research agenda connecting post-training methods to quantitative finance applications
  • Collaborate with quant researchers to define task distributions and evaluation frameworks
  • Unblock production systems dependent on post-training capabilities

Skills

PyTorch/JAX
Distributed training
RLHF/DPO knowledge
LLM alignment research
Experiment design
Budgeting compute

Education

BS or equivalent in STEM
MS preferred

Tools

DeepSpeed
FSDP

Job description

Two Sigma is advancing post-training with RLHF, DPO, and reward modeling to align LLMs with complex, multi-step financial workflows. You will define the research agenda, build scalable infrastructure, and guide evaluation frameworks for transforming quant research into production-grade AI capabilities.

The role combines training, fine-tuning, context management, and model evaluation, shaping both the post-training capability and the broader research direction of the team.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Post-Training Research Scientist
Post-Training Research Scientist

Two Sigma • New York (NY)

Hybrid
USD 165,000 - 300,000
Health insurance
401k match
Life & disability insurance
+5
AI Research Scientist (MS/PhD) — ML, LLMs & RL for Finance
AI Research Scientist (MS/PhD) — ML, LLMs & RL for Finance

Two Sigma • New York (NY)

On-site
LLM Post-Training Research Scientist (SFT & RLHF)
LLM Post-Training Research Scientist (SFT & RLHF)

Scale AI, Inc. • New York (NY)

On-site
USD 181,000 - 226,000
Health, dental & vision coverage
Retirement benefits
Learning and development stipend
+2
Post-Training ML Research Scientist (RLHF/SFT)
Post-Training ML Research Scientist (RLHF/SFT)

United States Digital Space LLC • New York (NY), San Francisco (CA)

On-site
USD 181,000 - 226,000
Health coverage
Equity
Generous PTO
+2
AI Research Intern: Deep Learning, LLMs & RL for Finance
AI Research Intern: Deep Learning, LLMs & RL for Finance

Two Sigma • New York (NY)

On-site
USD 228,000 - 250,000
AI Researcher — Finance Intelligence & RL Systems
AI Researcher — Finance Intelligence & RL Systems

ScOp Venture Capital LLC. • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
LLM Post-Training Scientist - SFT & RLHF
LLM Post-Training Scientist - SFT & RLHF

Scale AI, Inc. • San Francisco (CA)

On-site
USD 181,000 - 226,000
Health benefits
Retirement plan
Learning stipend
+2
LLM Post-Training Research Scientist (SFT/RLHF)
LLM Post-Training Research Scientist (SFT/RLHF)

Scale AI • New York (NY)

On-site
USD 181,000 - 226,000
Health, dental and vision coverage
Retirement benefits
Learning & development stipend
+2
Senior Research Scientist, Long-Horizon AI & RL (Equity)
Senior Research Scientist, Long-Horizon AI & RL (Equity)

techire ai • San Francisco (CA)

On-site
USD 400,000 - 450,000
Principal Research Scientist: AI Systems & RL in Production
Principal Research Scientist: AI Systems & RL in Production

Centific • East Palo Alto (CA)

On-site
USD 250,000 - 300,000