Research Scientist, Evaluations

Tenera, Inc

San Francisco (CA)

On-site

USD 120,000 - 180,000

Full time

24 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Tenera, Inc. seeks a Research Scientist in Evaluations to lead an ambitious, open-ended research program with substantial uncertainty. You will define questions, develop original approaches, and test ideas with scientific rigor, enjoying autonomy for independent work.

You will design experiments, build evaluation benchmarks, and manage data-labeling pipelines, documenting uncertainty and ground truth while driving results to publishable conclusions.

Qualifications

  • Track record of original empirical or methodological research in statistics, machine learning, behavioral science, psychometrics, computational social science, or a related field.
  • Exceptional creativity and execution in building experiments and solving problems.
  • Comfort with substantial uncertainty and sustained independent work; ability to publish results and revise conclusions.

Responsibilities

  • Define and pursue an ambitious research agenda in simulation evaluation.
  • Turn loosely defined questions into concrete research plans with milestones.
  • Design and implement evaluation methods grounded in observed behavior and data.
  • Build and operate data-labeling pipelines with guidelines and quality checks.
  • Design controlled comparisons, holdout datasets, and replication studies.
  • Study validity, reliability, and generalization across populations and contexts.
  • Own execution from hypothesis through reproducible results and scientific writing.

Skills

Python
Data analysis
Statistical inference
Experimental design
Scientific writing
Independent research

Tools

Annotation tools

Job description

Tenera helps teams make product decisions using simulated users. Understanding when those simulations can be trusted raises difficult questions about measurement, evidence, and generalization.

As a Research Scientist in Evaluations, you will own an ambitious, open-ended research agenda with substantial uncertainty. The questions, methods, and path to a useful result will not arrive fully defined. You will help define them, develop original approaches, and test your ideas with scientific rigor.

This role requires exceptional creativity and execution in equal measure. You will have significant autonomy and sustained time for independent research, with responsibility for turning uncertain ideas into concrete progress. You should be comfortable building experiments yourself, seeking critical feedback, and pursuing difficult questions even when early attempts fail.

You will have a dedicated research budget to collect data and build benchmarks, with the autonomy and responsibility to decide how those resources can best advance your research.

02 What You’ll Do
  • Define and pursue an ambitious research agenda in simulation evaluation. Decide which questions matter, challenge existing assumptions, and develop original approaches when established methods are insufficient.
  • Turn loosely defined questions into concrete research plans. Set milestones, prioritize experiments, and revise your direction as evidence emerges.
  • Design and implement evaluation methods grounded in observed behavior. Build the datasets, experimental tools, and analyses needed to test your ideas end to end.
  • Build and operate data-labeling pipelines, including annotation guidelines, annotator calibration, quality checks, and adjudication. Establish traceable ground truth and document uncertainty or disagreement in the labels.
  • Design controlled comparisons, holdout datasets, and replication studies. Account for sampling bias, data leakage, confounding, and uncertainty before drawing conclusions.
  • Study validity, reliability, and generalization across populations and contexts. Seek out counterexamples, investigate failures, and distinguish a promising result from a defensible conclusion.
  • Own execution from the first hypothesis through reproducible results and clear scientific writing. Seek targeted criticism from teammates while remaining responsible for research direction and progress.
03 What We’re Looking For
  • A track record of original empirical or methodological research in statistics, machine learning, behavioral science, psychometrics, computational social science, or a related field. Demonstrate your contribution through publications, research reports, or substantial independent projects.
  • Exceptional creativity. You can identify questions others have overlooked, connect ideas across disciplines, and devise useful methods when familiar approaches do not work.
  • Exceptional execution. You turn ambitious ideas into working experiments, resolve practical obstacles, and carry difficult projects through to concrete, verifiable results.
  • Comfort with substantial uncertainty and sustained independent work. You can make progress without a detailed brief, tolerate inconclusive experiments, and change course without losing sight of the research goal.
  • Strong foundations in statistical inference and experimental design. You can reason about measurement validity, statistical power, uncertainty, selection bias, and the limits of causal claims.
  • Prior experience building evaluation benchmarks end to end is required, from defining the task and sampling data to establishing ground truth, designing scoring methods, and validating the benchmark.
  • Prior experience building and operating data-labeling pipelines is required. You have written annotation guidelines, calibrated annotators, measured agreement, resolved ambiguous labels, and implemented quality controls and dataset versioning.
  • Strong Python and data analysis skills. You can implement research methods, work with real-world datasets, and build reproducible experiments that others can inspect.
  • Clear scientific writing and intellectual honesty. You seek criticism, revise your conclusions, and report null results and limitations with the same care as positive findings.
  • Humble. You are quick to learn from customers, teammates, and evidence instead of defending your first answer.
  • Low ego. You care more about the best idea winning than being personally right.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Scientist, Evaluations - Simulation Benchmarks
Research Scientist, Evaluations - Simulation Benchmarks

Tenera, Inc • San Francisco (CA)

On-site
USD 120,000 - 180,000
Research Scientist, Modeling
Research Scientist, Modeling

Tenera, Inc • San Francisco (CA)

On-site
USD 150,000 - 210,000
Equity
Relocation assistance
Research Scientist
Research Scientist

Anyone AI Inc. • Northern (KY)

On-site
USD 110,000 - 160,000
Member of Technical Staff — Research Engineering, Evaluation
Member of Technical Staff — Research Engineering, Evaluation

Causal Labs • San Francisco (CA)

On-site
USD 180,000 - 240,000
Head of Research
Head of Research

talp • San Francisco (CA)

On-site
USD 180,000 - 320,000
Member of Technical Staff, Deployment Strategist
Member of Technical Staff, Deployment Strategist

Tenera, Inc • San Francisco (CA)

On-site
USD 120,000 - 180,000
Member of Technical Staff — Research Engineering, Evaluation
Member of Technical Staff — Research Engineering, Evaluation

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 200,000
Research Engineer, Evals
Research Engineer, Evals

Variance • San Francisco (CA)

On-site
USD 170,000 - 230,000
Competitive salary
Platinum-level medical, dental, and vision insurance
Unlimited PTO
+6
Applied Research — Evaluations & Data
Applied Research — Evaluations & Data

Human Intuition Inc. • New York (NY)

On-site
USD 120,000 - 180,000
Research Scientist
Research Scientist

Tessera Labs • San Jose (CA)

On-site
USD 140,000 - 210,000