Post-Training Research Scientist

Two Sigma

New York (NY)

Hybrid

USD 165,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401k match
Life & disability insurance
Gym access
Tuition reimbursement
Generous vacation & sick days
Hybrid work policy
Home office budget

Job summary

Two Sigma is advancing post-training with RLHF, DPO, and reward modeling to align LLMs with complex, multi-step financial workflows. You will define the research agenda, build scalable infrastructure, and guide evaluation frameworks for transforming quant research into production-grade AI capabilities.

The role combines training, fine-tuning, context management, and model evaluation, shaping both the post-training capability and the broader research direction of the team.

Qualifications

  • BS or equivalent in Science, Technology, Engineering or Math (an MS is a plus).
  • Minimum 1 year of experience required; 1-10 years preferred at a frontier AI lab (OpenAI, Anthropic, DeepMind, Meta FAIR, or equivalent).
  • Shipped post-training systems in production: RLHF, DPO, RLAIF, or related methods.
  • Deep understanding of distributed training infrastructure: multi-node GPU clusters, training stability, checkpointing.
  • Track record managing large-scale compute: budgeting, experiment design, ablations.
  • Publications or demonstrated expertise in alignment, preference learning, or reward modeling.
  • Hands-on implementation skills: PyTorch/JAX, distributed frameworks (DeepSpeed, FSDP, etc.)

Responsibilities

  • Lead post-training efforts for LLMs applied to financial time series and quantitative reasoning
  • Design and execute RLHF, DPO, and related alignment methods at scale, including deployment of substantial compute budgets (O($100mm))
  • Build infrastructure for preference data collection, reward modeling, and policy optimization on financial datasets
  • Drive research agenda connecting post-training methods to quantitative finance applications
  • Collaborate with quant researchers to define task distributions and evaluation frameworks
  • Unblock production systems dependent on post-training capabilities

Skills

PyTorch/JAX
Distributed training
RLHF/DPO knowledge
LLM alignment research
Experiment design
Budgeting compute

Education

BS or equivalent in STEM
MS preferred

Tools

DeepSpeed
FSDP

Job description

Position Summary

Two Sigma is a leading quantitative investment management and trading firm. The company applies a scientific approach to investing, combining cutting-edge technology, artificial intelligence, data science, and quantitative research with rigorous human inquiry to capitalize on market opportunities and deliver alpha for investors.

Our team of engineers, quantitative researchers and data scientists looks beyond the traditional to test hypotheses and develop creative solutions to some of the world’s most complex economic problems.

We are applying large language models and transformer-based architectures to problems where ground truth is delayed, noisy, and non-stationary. Our systems generate code, run experiments, and iterate autonomously, and we are looking to go beyond supervised fine-tuning.

We are hiring a Post-Training Research Scientist to build RLHF, DPO, and reward modeling capabilities from the ground up. This is a greenfield role: you will define the infrastructure, research agenda, and evaluation frameworks for aligning LLMs to sophisticated, multi-step workflows in a domain where the reward signal is fundamentally different from existing research on human preference or deterministic task completion.

This hire will help own methodology across training, fine-tuning, context management, and model evaluation. You will shape not only the post-training capability but the broader research direction of the team.

You Will Take On The Following Responsibilities
  • Lead post-training efforts for LLMs applied to financial time series and quantitative reasoning
  • Design and execute RLHF, DPO, and related alignment methods at scale, including deployment of substantial compute budgets (O($100mm))
  • Build infrastructure for preference data collection, reward modeling, and policy optimization on financial datasets
  • Drive research agenda connecting post-training methods to quantitative finance applications
  • Collaborate with quant researchers to define task distributions and evaluation frameworks
  • Unblock production systems dependent on post-training capabilities
You Should Possess The Following Qualifications
  • BS or equivalent work experience in Science, Technology, Engineering or Math (an MS is a plus).
  • Minimum 1 year of experience required; 1-10 years of experience preferred (ideally 1-5 years) at a frontier AI lab (OpenAI, Anthropic, DeepMind, Meta FAIR, or equivalent)
  • Shipped post-training systems in production: RLHF, DPO, RLAIF, or related methods
  • Deep understanding of distributed training infrastructure: multi-node GPU clusters, training stability, checkpointing
  • Track record managing large-scale compute: budgeting, experiment design, ablations
  • Publications or demonstrated expertise in alignment, preference learning, or reward modeling
  • Hands-on implementation skills: PyTorch/JAX, distributed frameworks (DeepSpeed, FSDP, etc.)
You Will Enjoy The Following Benefits
  • Core Benefits: Fully paid medical and dental insurance premiums for employees and dependents, competitive 401k match, employer-paid life & disability insurance
  • Perks: Onsite gyms with laundry service, wellness activities, casual dress, snacks, game rooms
  • Learning: Tuition reimbursement, conference and training sponsorship
  • Time Off: Generous vacation and unlimited sick days, competitive paid caregiver leaves
  • Hybrid Work Policy: Flexible in-office days with budget for home office setup

The base pay for this role will be between $165,000 and $300,000. This role may also be eligible for other forms of compensation and benefits, such as a discretionary bonus, health, dental and other wellness plans and 401(k) contributions. Discretionary bonus can be a significant portion of total compensation. Actual compensation for successful candidates will be carefully determined based on a number of factors, including their skills, qualifications and experience.

We are proud to be an equal opportunity workplace. We do not discriminate based upon race, religion, color, national origin, sex, sexual orientation, gender identity/expression, age, status as a protected veteran, status as an individual with a disability, or any other applicable legally protected characteristics.

Two Sigma is committed to providing reasonable accommodations to qualified individuals in accordance with applicable federal, state, and local laws.

If you believe you need an accommodation, please visit our website for additional information.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Research Scientist - Campus Full-Time
AI Research Scientist - Campus Full-Time

Two Sigma • New York (NY)

On-site
USD 200,000 - 220,000
Onsite gyms with laundry service
Wellness activities
Casual dress
+4
Quantitative Researcher - Experienced Hire
Quantitative Researcher - Experienced Hire

Two Sigma • New York (NY)

Hybrid
USD 165,000 - 325,000
Fully paid medical and dental insurance
Competitive 401k match
Onsite gyms with laundry service
+2
Quantitative Researcher - Full-Time Campus Hire
Quantitative Researcher - Full-Time Campus Hire

Two Sigma • New York (NY)

Hybrid
USD 200,000 - 220,000
Fully paid medical and dental insurance
Competitive 401k match
Tuition reimbursement
+2
Research Scientist, Post-Training
Research Scientist, Post-Training

David Joseph & Company • San Francisco (CA)

On-site
USD 150,000 - 450,000
Research, Post-Training Data
Research, Post-Training Data

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research, Post-Training
Research, Post-Training

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health benefits
Dental and vision benefits
Unlimited PTO
+2
Quantitative Researcher: Machine Learning
Quantitative Researcher: Machine Learning

Quant Blueprint LLC • New York (NY)

Hybrid
USD 165,000 - 325,000
Fully paid medical and dental insurance premiums
Onsite gyms and wellness activities
Tuition reimbursement and training sponsorship
+2
Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Competitive cash and equity compensation (>90th percentile)
Ownership and autonomy
Health, vision, dental benefits
+4
Quantitative Software Engineer: Generative AI
Quantitative Software Engineer: Generative AI

Two Sigma • New York (NY)

Hybrid
USD 165,000 - 300,000
Fully paid medical and dental insurance
Tuition reimbursement
Flexible in-office days
+1
Research, Post-Training Data
Research, Post-Training Data

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3