Get more replies from employers
Send a job-specific resume in minutes.
Preference Model is seeking a Research Engineer or Research Scientist located in Seattle, WA, to innovate in the field of self-directed learning for models. The role involves training and evaluating models in custom RL settings and optimizing ML infrastructure.
Ideal candidates will have experience with LLM post-training pipelines and be proficient in Python and modern RL frameworks. Preference Model offers a competitive compensation package, the opportunity to work with top engineers, and a challenging startup environment.
About Us
Preference Model is building automated ML research engineering. Existing frontier models are brittle when applied to real-world ML tasks. The present bottleneck is the lack of high-quality RL training environments. Our first step is to build RL environments that reflect real-world complexity, with diverse tasks and robust reward functions. Our founding team has previous experience on Anthropic’s data team building data infrastructure, and datasets behind Claude. We are partnering with leading AI labs to push AI closer to achieving its transformative potential.
Models of the future will be able to train themselves on tasks that they are not good at. We are interested in investigating how far we can push the boundaries of self-directed learning. We are looking for Research Engineers or Research Scientists to push the frontier of post-training on large language models in a role that blends research and engineering, requiring you to implement novel approaches and shape research directions.
Candidates don't need a PhD or extensive publications. Some of the best researchers have no formal ML training and gained experience building industry products. We believe adaptability combined with exceptional communication and collaboration skills are the most important ingredients for successful startup research.
We value diverse perspectives and experiences. If you're excited about this role but don't check every box, we still encourage you to apply.
Compensation Range: $200K - $350K