Get more replies from employers
Send a job-specific resume in minutes.
Preference Model is seeking Research Engineers or Research Scientists to advance self-directed learning in AI. The role involves training and evaluating models within proprietary RL environments and optimizing ML infrastructure.
Candidates will benefit from competitive compensation, ownership in a fast-paced startup, and support for personal growth. Experience in Python, PyTorch, or JAX, along with RL training frameworks, is desired. The opportunity is a unique blend of research and engineering within a dynamic team.
Preference Model is building automated ML research engineering. Existing frontier models are brittle when applied to real-world ML tasks. The present bottleneck is the lack of high-quality RL training environments. Our first step is to build RL environments that reflect real-world complexity, with diverse tasks and robust reward functions.
Our founding team has previous experience on Anthropic’s data team building data infrastructure, and datasets behind Claude. We are partnering with leading AI labs to push AI closer to achieving its transformative potential.
Models of the future will be able to train themselves on tasks that they are not good at. We are interested in investigating how far we can push the boundaries of self‑directed learning. We are looking for Research Engineers or Research Scientists to push the frontier of post‑training on large language models in a role that blends research and engineering, requiring you to implement novel approaches and shape research directions.
Candidates don't need a PhD or extensive publications. Some of the best researchers have no formal ML training and gained experience building industry products. We believe adaptability combined with exceptional communication and collaboration skills are the most important ingredients for successful startup research.