Get more replies from employers
Send a job-specific resume in minutes.
Aaru in New York City is seeking an Evaluation Research Manager to lead a focused team of researchers and engineers. You will translate the evaluation charter into a portfolio of studies, datasets, and reusable infrastructure, ensuring the quality, pace, and usefulness of the team's work.
You will design studies, write analysis code, review statistical choices, and contribute to the hardest evaluations while hiring and coaching researchers, and collaborating with cross-functional teams to
Aaru builds simulations of human behavior. Each simulation contains a population of AI agents, each representing a person who could plausibly exist in the real world and capable of making decisions within a modeled environment. Companies and institutions use these simulations to test consequential choices before committing—from product launches and pricing decisions to strategic communications and policy changes. Building a useful simulation requires more than generating plausible text. Populations must represent real people and groups; predictions must be calibrated; simulations must remain coherent as conditions change; and the product must make the resulting evidence legible enough to support real decisions. We are a small, in-person team in New York. We work with urgency, high ownership, and intellectual honesty. We expect people to surface inconvenient evidence, change their minds quickly, and carry important work all the way to a result.
Evaluation Research determines whether Aaru's populations, predictions, and end-to-end simulations correspond closely enough to the real world to support consequential decisions. The team defines what should be measured, develops the methods for measuring it, and produces the evidence Aaru uses to improve its systems and describe their capabilities. The function builds both rails and carts. Rails are reusable evaluation infrastructure: datasets, harnesses, libraries, experiment standards, leaderboards, reporting systems, and ways to translate technical evidence into decisions. Carts are the specific evaluations that run on those rails: a historical backtest, a prospective forecast study, a population-coherence test, a reproduction of an observed behavioral pattern, or an end-to-end comparison with a resolved outcome. Evaluation Research is not conventional QA and it is not an internal approval service. It is an independent research function that works closely with the teams building Aaru's systems while preserving the ability to reach and communicate inconvenient conclusions.
As Evaluation Research Manager, you will lead a focused team of Evaluation Researchers and research engineers. You will translate Aaru's evaluation charter into a coherent portfolio of studies, datasets, and shared infrastructure, and you will be accountable for the quality, pace, and usefulness of the team's work. Managers at Aaru remain researchers. You will design studies, write analysis code, inspect individual failures, review statistical and measurement choices, and directly contribute to the hardest evaluations. You will also set priorities, hire exceptional people, develop the team, provide candid feedback, and create the operating mechanisms that keep protected evidence independent while making diagnostic evidence available quickly. You will work across Population Research, Prediction Research, Simulation Engineering, Product Engineering, Research Product, Deployment, and company leadership. The job requires both scientific independence and practical judgment: an evaluation must be rigorous enough to support a real claim and diagnostic enough to help a team improve the system.
Build, lead, and develop a high-performing team of Evaluation Researchers and research engineers.
Turn broad questions about realism, accuracy, calibration, usefulness, and decision quality into measurable constructs, decisive experiments, and explicit decision criteria.
Set a focused evaluation agenda across population construction, predictive systems, individual agent behavior, group dynamics, and end-to-end simulations.
Decide which evaluation infrastructure should become a reusable organizational rail and which questions require a purpose-built study.
Establish standards for baselines, temporal holdouts, prospective testing, contamination control, statistical power, uncertainty, subgroup analysis, and reproducibility.
Build tests of population quality that assess individual coherence, joint and conditional distributions, representation of rare but plausible profiles, and whether a profile induces behavior consistent with the person it represents.
Evaluate forecasts and other predictive outputs using calibration, proper scoring rules, ranking quality, selective prediction, temporal validity, subgroup performance, and the real cost of different errors.
Compare simulations with transactions, product usage, behavioral traces, operational outcomes, market movements, resolved events, and longitudinal decisions.
Design end-to-end studies that determine whether improvements to a component actually improve the decision-relevant output customers receive.
Find failures hidden by aggregate metrics, especially those concentrated in important subgroups, rare cases, changing environments, or ambiguous ground truth.
Create diagnostic evaluations that help researchers localize why a system failed and distinguish a real general improvement from benchmark-specific optimization.
Partner with Simulation Engineering to make evaluations repeatable, versioned, scalable, and integrated into development and release workflows without compromising protected holdouts.
Convert production incidents, customer surprises, and deployment failures into durable test cases and better measurement methods.
Review evidence used in product, customer, or public claims and ensure that conclusions are reproducible, appropriately scoped, and honest about uncertainty and limits.
Communicate negative, null, and inconclusive results with the same precision and urgency as positive findings.
Recruit exceptional researchers, set clear expectations, provide direct feedback, develop independent research judgment, and address performance problems early.
You might be responsible for situations such as:
Useful evaluation measures something consequential and helps the company learn. We compare systems with strong alternatives, protect held-out data, quantify uncertainty, and separate exploratory findings from evidence used to support a claim. When ground truth is imperfect, the quality and limits of the outcome data are part of the research problem rather than a footnote. Evaluation should accelerate research by giving teams clear, diagnostic signals about what improved, what did not, and why. It should also make Aaru more trustworthy by exposing failures early and keeping product and external claims aligned with the available evidence. Management in this function requires independence without isolation. The team must understand the systems deeply enough to measure them well, collaborate closely enough to make the results useful, and remain willing to conclude that an attractive idea did not work.
This role is based in New York City. Aaru is an in-person company, working five days a week in the office. Candidates should be located in the New York metropolitan area or open to relocation.