Research Intern

Kindredventures

San Francisco (CA)

On-site

USD 34,440,000 - 55,104,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

In-office perks
Daily lunch
Housing stipend

Job summary

Scorecard is seeking a Research Intern in San Francisco to own one research question end-to-end within simuser frameworks for AI agent simulations. You’ll design experiments, build prototypes, evaluate outcomes, and publish findings with a dedicated mentor.

Internships run 10 to 16 weeks with flexible start dates; you’ll work with real data (human work transcripts), ship your prototypes into the framework, and contribute to public benchmarks or blog posts.

Qualifications

  • You have conducted rigorous experiments and can design studies.
  • You have experience with language models and related experiments (evals, scaffolding, data pipelines).
  • You can articulate what makes results trustworthy.

Responsibilities

  • Own a research question end to end: literature, design, prototype, evaluate, write-up.
  • Work with real data (transcripts) and assess what it can reveal.
  • Ship prototypes into Scorecard's framework and platform.
  • Publish results as a paper or public benchmark.

Skills

Research experience
LLM experience
Agent behavior sense
Ownership

Education

MS or PhD student or comparable research

Tools

Python

Job description

Research Intern
About Scorecard

We're a small team within Scorecard, the simulation platform for AI agents, and we work closely with our customers to build multi-user simulations.

The role

This team focuses on creating simulations for agents interacting with multiple users, which means we need simulated users ("simusers") to interact with the agent. Currently, engineers simulate users using vibecoded prompts like "act frustrated," resulting in unrealistic users that corrupt evaluation results.

To fix this, we’re building a simuser framework and orchestration layer. It allows defining high-fidelity simusers and realistically managing a group of simusers that interact with the agent-under-test in a simulation.

We're hiring a research intern to own one question in this space end-to-end and publish what they find, paired with a dedicated mentor from the team throughout. Internships run 10 to 16 weeks with flexible start dates throughout the year.

What you'll do
  • Own a research question end to end. Literature, experiment design, prototype, evaluation, write-up. We scope the project to your strengths when you arrive. Example questions: how do we measure that a simuser behaves like the specific person it was built from? What can we recover about someone's goals, knowledge, and style from their work transcripts? How can we avoid hallucination and sim drift in long-horizon simulations?
  • Work with real data. Our grounding data is human work transcripts. Expect to spend time cleaning it and to develop opinions about what it can and can't tell you.
  • Ship into the framework. Your prototypes land in the codebase, and your experiments run against the same system Scorecard's platform and partners use.
  • Publish. You'll be an author on the work you ship. We expect strong results to become a paper or public benchmark, and we write up intern projects on our blog.
What we're looking for
  • Research experience. You're an MS or PhD student, or have done comparable research outside a degree program. You've turned vague questions into rigorous experiments and know what makes a result trustworthy.
  • LLM experience. You've built with language models before, e.g. evals, agent scaffolding, or data pipelines, and can run experiments with Python.
  • A feel for agent behavior. You notice when a simulated user sounds off, and you can explain why and check whether a fix worked.
  • Hands-on ownership. The framework is new and the team is small. You're comfortable making design decisions and experimenting with incomplete information.

These aren't hard requirements and you don't need publications. If you've done solid research and this problem interests you, apply!

Nice to have
  • Published research: papers, benchmarks, datasets, or public experiments.
  • Prior work on user simulation: simulated users for dialogue or agent benchmarks, digital twins, role-play systems, or character consistency.
  • Evaluation work: LLM-as-judge design, grader calibration, or human annotation pipelines.
  • RL experience: building RL environments for LLMs, reward design, or firsthand encounters with reward hacking.
  • Behavioral or social-science research methods: survey design, test-retest protocols, inter-rater reliability.
Compensation and benefits
  • In-office perks and daily lunch
  • Competitive hourly rate + housing stipend
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Research Engineer
Research Engineer

Kindredventures • San Francisco (CA)

On-site
USD 175,000 - 250,000
In-office perks and daily lunch
Medical, dental, and vision benefits
Unlimited PTO
+1
Research Intern: Simuser Framework & Evaluation
Research Intern: Simuser Framework & Evaluation

Kindredventures • San Francisco (CA)

On-site
USD 34,440,000 - 55,104,000
In-office perks
Daily lunch
Housing stipend
Researcher [33023]
Researcher [33023]

Stealth Startup • New York (NY)

On-site
USD 120,000 - 160,000
Medical coverage
Vision coverage
Dental coverage
+2
Research Intern
Research Intern

Quadrillion Labs • New York (NY)

On-site
USD 300,000 - 500,000
Medical, dental, and vision insurance
Lunch and dinner covered
Other varied stipends
Research Infrastructure - Member of Technical Staff
Research Infrastructure - Member of Technical Staff

Simile • San Francisco (CA)

On-site
USD 200,000 - 400,000
Equity grants
Health, dental, vision
Flexible time off
Research Intern
Research Intern

jobr.pro • New York (NY)

On-site
USD 30,000 - 40,000
Medical, dental, and vision insurance
Lunch and dinner covered
Generous stipends
Research Intern
Research Intern

NeoCognition Inc. • Palo Alto (CA)

On-site
USD 60,000 - 80,000
Research Intern: Interpretability & Reliability (Summer 2027)
Research Intern: Interpretability & Reliability (Summer 2027)

CTGT • San Francisco (CA)

On-site
USD 52,000 - 68,000
Research Intern
Research Intern

NeoCognition • Palo Alto (CA)

On-site
Research - Member of Technical Staff
Research - Member of Technical Staff

Simile • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health & Wellness
Equity
Flexible time off