Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Techire AI is seeking a Senior Research, Evals to rethink how model performance is measured across real-world interactions. You will develop evaluation frameworks that go beyond static benchmarks and build systems that integrate evaluations into model development.
The role is hands-on, bridging model development and evaluation, with responsibilities to design studies, pipelines and metrics for reasoning, memory, and interaction quality.
How do you actually evaluate an AI model when accuracy on a benchmark only tells part of the story?
A well-funded AI company developing next-generation foundation models is looking for a Senior Research, Evals to rethink how model performance is measured across real-world interactions.
You’ll work on evaluation frameworks that go beyond static benchmarks, looking at how models reason, remember, adapt and interact over time.
This is a hands‑on research role sitting between model development and evaluation. You’ll help define what should be measured, work out how to measure it, then build the systems that bring those evaluations directly into model development.
Experience with alignment or model behaviour research would be useful, but isn’t essential.
You’ll work closely with researchers building new foundation models, helping understand whether changes are creating meaningful improvements rather than simply moving benchmark scores.
Base salary is $200k-$350k DOE + generous equity.
Based in San Francisco, New York or London, working hybrid.