Stand out for this role — generate a tailored resume and cover letter in about a minute.
techire ai is seeking a researcher to push the frontiers of long-horizon reasoning and RL, building systems that outperform current baselines. You will tackle problems without standard solutions and shape memory and reasoning components for real-world impact.
Ideal candidates hold a PhD with top-conference publications and have hands-on post-training RL experience, including RLHF or reward modelling, plus a track record in open-ended research.
Rip up the playbook and step into uncharted territory.
If you've been building long-horizon multi-agent systems and pushing the boundaries of AI research, this is the kind of role where curiosity and ambition meet real execution, exploring truly novel problems at the frontier of what's currently possible.
You will work on systems designed to outperform the current state of the art, tackling problems that don't yet have standardised solutions across RL, long-horizon reasoning, LLM post-training for non-myopic objectives, environment and feedback design.
Whether you're early-career PhD or highly experienced, what matters most is your ability to push novel ideas into working systems, execute your knowledge across reasoning, RL and memory to make real-world impact.
This is a small, ambitious team operating where few others are, building and executing quickly in areas such as computational R&D science. This is your opportunity to shape the systems that generate and validate new discovery in environment primed for success.
San Francisco
$400k base 0.5–1%+ equity Negotiable DOE
All applicants will receive a response.