Research Scientist, Evaluations - Simulation Benchmarks

Tenera, Inc

San Francisco (CA)

On-site

USD 120,000 - 180,000

Full time

31 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Tenera, Inc. seeks a Research Scientist in Evaluations to lead an ambitious, open-ended research program with substantial uncertainty. You will define questions, develop original approaches, and test ideas with scientific rigor, enjoying autonomy for independent work.

You will design experiments, build evaluation benchmarks, and manage data-labeling pipelines, documenting uncertainty and ground truth while driving results to publishable conclusions.

Qualifications

  • Track record of original empirical or methodological research in statistics, machine learning, behavioral science, psychometrics, computational social science, or a related field.
  • Exceptional creativity and execution in building experiments and solving problems.
  • Comfort with substantial uncertainty and sustained independent work; ability to publish results and revise conclusions.

Responsibilities

  • Define and pursue an ambitious research agenda in simulation evaluation.
  • Turn loosely defined questions into concrete research plans with milestones.
  • Design and implement evaluation methods grounded in observed behavior and data.
  • Build and operate data-labeling pipelines with guidelines and quality checks.
  • Design controlled comparisons, holdout datasets, and replication studies.
  • Study validity, reliability, and generalization across populations and contexts.
  • Own execution from hypothesis through reproducible results and scientific writing.

Skills

Python
Data analysis
Statistical inference
Experimental design
Scientific writing
Independent research

Tools

Annotation tools

Job description

Tenera, Inc. seeks a Research Scientist in Evaluations to lead an ambitious, open-ended research program with substantial uncertainty. You will define questions, develop original approaches, and test ideas with scientific rigor, enjoying autonomy for independent work.

You will design experiments, build evaluation benchmarks, and manage data-labeling pipelines, documenting uncertainty and ground truth while driving results to publishable conclusions.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Scientist, Evaluations
Research Scientist, Evaluations

Tenera, Inc • San Francisco (CA)

On-site
USD 120,000 - 180,000
Research Evaluations Scientist - Post-Training & Benchmarks
Research Evaluations Scientist - Post-Training & Benchmarks

Thinking Machines Lab • San Francisco (CA)

On-site
USD 350,000 - 475,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Staff Engineer, Evaluation Infrastructure
Staff Engineer, Evaluation Infrastructure

Simile • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health & Wellness
Equity
Flexible time off
AI Evaluations Engineer — Benchmarking Frontiers
AI Evaluations Engineer — Benchmarking Frontiers

Meta • Menlo Park (CA)

On-site
USD 180,000 - 240,000
Frontier Benchmarking Research Engineer
Frontier Benchmarking Research Engineer

Refresh • San Francisco (CA)

On-site
USD 120,000 - 190,000
Research Engineer — RL Evaluation & Benchmarks
Research Engineer — RL Evaluation & Benchmarks

Invisible Technologies Inc. • New York (NY)

Hybrid
USD 150,000 - 210,000
Bonuses and equity
Hybrid work environment
Research Scientist, Modeling
Research Scientist, Modeling

Tenera, Inc • San Francisco (CA)

On-site
USD 150,000 - 210,000
Equity
Relocation assistance
Remote Senior AI/ML Benchmarking & Evaluation Engineer
Remote Senior AI/ML Benchmarking & Evaluation Engineer

OpenTeams • Northern (KY)

Hybrid
USD 145,000 - 250,000
LLM Evaluation Engineer - Benchmarks & Failure Analysis
LLM Evaluation Engineer - Benchmarks & Failure Analysis

Nous Research • United States

Remote
USD 130,000 - 170,000
Staff Engineer, Model Evaluations & Behavioral Metrics
Staff Engineer, Model Evaluations & Behavioral Metrics

Simile • New York (NY), Northern (KY)

Hybrid
USD 200,000 - 400,000
Health & Wellness
Time Off
Equity grants