Turn this role into an interview — a resume and cover letter built around what this employer wants.
Two Sigma is advancing post-training with RLHF, DPO, and reward modeling to align LLMs with complex, multi-step financial workflows. You will define the research agenda, build scalable infrastructure, and guide evaluation frameworks for transforming quant research into production-grade AI capabilities.
The role combines training, fine-tuning, context management, and model evaluation, shaping both the post-training capability and the broader research direction of the team.
Two Sigma is advancing post-training with RLHF, DPO, and reward modeling to align LLMs with complex, multi-step financial workflows. You will define the research agenda, build scalable infrastructure, and guide evaluation frameworks for transforming quant research into production-grade AI capabilities.
The role combines training, fine-tuning, context management, and model evaluation, shaping both the post-training capability and the broader research direction of the team.