Stand out for this role — generate a tailored resume and cover letter in about a minute.
Alexander Chapman is seeking a Forward Deployed ML Engineer, Benchmarks & Evaluations to help build the infrastructure behind AI evaluations and data benchmarks. You will be one of the first engineers in this high-ownership, fast-moving team, collaborating with the GM, researchers and customers to deploy repeatable evaluation products.
The role focuses on building LLM benchmarks, data pipelines, sandboxed evaluation environments, and close collaboration with customers to shape evaluation
Building the infrastructure for AI training data and evaluations, helping frontier AI teams securely access the data they need to improve models. The team is early-stage, high-ownership, and focused on moving quickly while solving difficult technical problems.
You will be one of the first engineers focused on the Benchmarks & Evaluations business, working closely with the GM, researchers, and customers to build the infrastructure behind next-generation AI evaluations.
experience with LLM evals/benchmarks, human data pipelines, frontier AI labs, agentic systems, RL environments, or early-stage startups.