Get more replies from employers
Send a job-specific resume in minutes.
Artificial Analysis is seeking a Member of Technical Staff to design frontier evaluations for language models and publish results used by AI labs and enterprises. You will build datasets, scoring systems, and evaluation infrastructure applicable across major models released by labs worldwide.
Based in San Francisco or flexible to other major cities, this role offers equity and the opportunity to shape how industry measures frontier AI capability while collaborating with leading researchers and
Location: San Francisco (preferred), Sydney, Melbourne, Brisbane
Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand AI capabilities and make critical decisions about their AI strategies. We are the go-to authority for understanding AI, from AI labs and enterprises to media, investors, and policymakers. Our benchmarks don’t just measure the cutting edge of AI, they are actively shaping the frontier.
Our benchmarks and analysis are trusted by hundreds of thousands of users and are the go-to reference for leading AI labs including OpenAI, Google, Meta, NVIDIA and Anthropic, and major publications including the Wall Street Journal, Bloomberg, the Financial Times and The Economist.
We are a team of 40+, on track to double by end of year, backed by Nat Friedman (GitHub, Meta), Daniel Gross (SSI, Meta), Andrew Ng (Google Brain, DeepLearning.ai, Amazon), Adam D’Angelo (Quora, Poe, OpenAI), Clem Delangue (Hugging Face) and other industry leaders.
Language model evaluation is the sharpest question in AI: what can these systems actually do? Our answers, from the Artificial Analysis Intelligence Index to AA-Omniscience, AA-Briefcase and our coding agent evaluations, are the reference the industry uses. We’re hiring Members of Technical Staff to build the next generation of them.
This is a role for people who want to build frontier benchmarks: designing evaluations that stay ahead of frontier capabilities, constructing datasets that resist contamination, and measuring what everyone else has not yet worked out how to measure. You will run your work across every major model as it releases and publish results the whole industry reads.
The center of the role is building. Analysis and lab collaboration wrap around the evaluation work, with our commercial team owning client relationships day to day.
You have deep, hands-on experience evaluating language models and strong opinions about why most benchmarks fail.
Backgrounds include: evaluation and benchmarking teams at AI labs; research or engineering roles at evaluation-focused organizations; ML engineers who have built evaluation harnesses and datasets in production; or academic researchers in NLP and ML evaluation with a strong record of published work.
1