Get more replies from employers
Send a job-specific resume in minutes.
Artificial Analysis, based in San Francisco, is seeking a Forward Deployed Engineer to own the day‑to‑day operation of our language model benchmarking stack. You’ll onboard new models to our evaluation pipeline, run benchmarks, and serve as the primary technical contact for AI lab customers: explaining results, answering methodology questions, and resolving endpoint issues in real time.
This deployed role sits at the interface between our platform and sophisticated customers, requiring strong
Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand AI capabilities and make critical decisions about their AI strategies. We are the go-to authority for understanding AI, from AI labs and enterprises to media, investors, and policymakers. Our benchmarks don’t just measure the cutting edge of AI, they are actively shaping the frontier.
Our benchmarks and analysis are trusted by hundreds of thousands of users and are the go-to reference for leading AI labs including OpenAI, Google, Meta, NVIDIA and Anthropic, and major publications including the Wall Street Journal, Bloomberg, the Financial Times and The Economist.
We are a team of 40+, on track to double by end of year, backed by Nat Friedman (GitHub, Meta), Daniel Gross (SSI, Meta), Andrew Ng (Google Brain, DeepLearning.ai, Amazon), Adam D’Angelo (Quora, Poe, OpenAI), Clem Delangue (Hugging Face) and other industry leaders.
Artificial Analysis maintains one of the most comprehensive language model benchmarking suites in the industry, evaluating frontier models across quality, speed, and pricing for the AI labs and enterprises that rely on our data.
We’re hiring a Forward Deployed Engineer to own the day‑to‑day operation of our language model benchmarking stack and act as the technical face of Artificial Analysis to the industry’s most important labs. You’ll onboard new models to our evaluation pipeline, run and debug benchmarks, and serve as the primary technical point of contact for AI lab customers: explaining results, fielding methodology questions, and resolving endpoint issues in real time.
This is a deployed role in the truest sense: you sit at the interface between our platform and our most sophisticated customers, accountable for both sides working. It is about running a sophisticated stack exceptionally well, consistently and reliably, while being the trusted engineer our customers ask for by name.
Required:
Nice to have (not required):
1