Research Scientist

Heyaristotle

San Francisco (CA)

Hybrid

USD 85,000 - 120,000

Part time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Heyaristotle in San Francisco is looking for a talented researcher to work on improving AI tutoring systems. You will shape experiments, evaluate tutoring quality, and turn pedagogical principles into actionable measures.

The role starts as a summer contract with potential for continuation. Ideal candidates have a PhD or equivalent experience in CS, ML, or related fields. Responsibilities include planning research projects and building simulated student models.

Qualifications

  • Graduate research experience in CS, ML, NLP, or a related field.
  • Hands-on experience with LLM evaluation, post-training or fine-tuning.
  • Familiarity with education research is a strong plus.

Responsibilities

  • Scope and plan research projects with the team.
  • Design and validate measures of multi-turn tutoring quality.
  • Build simulated students grounded in real student data.

Skills

LLM evaluation
agent systems
dataset curation

Education

Graduate research experience (PhD student, recent PhD)

Job description

Run the experiments and build the evaluations that improve how our production AI tutor teaches real students.
About Aristotle

At Aristotle, we are building the world's first AI tutor — giving every student access to an expert, personal guide. Aristotle provides highly personalized tutoring, complete with memory, personality, consistency, and genuine care.

The Role

Research at Aristotle works on the open problems in building a realtime AI tutor. The biggest one right now is evaluation. Tutoring is multi-turn and adaptive, there is no verifiable reward signal, and existing benchmarks test single turns in toy scenarios. We are building the evaluations that make tutoring quality measurable, including simulated students realistic enough to test tutors against, and using them to improve our tutor in production.

You will own this work end to end: forming hypotheses, running experiments on real production sessions, and shipping findings into the product. You will work directly with one of the co-founders leading research, and with data that few groups have: full multimodal tutoring sessions at scale, with voice, whiteboard state, and student history.

This starts as a summer contract, with the intent to continue if it goes well. Full‑time in SF is preferred; we are flexible on hours and location for the right person. We support publishing where we can, and we will tell you the constraints up front.

Responsibilities
  • Scope and plan research projects with the team.
  • Design and validate measures of multi‑turn tutoring quality, grounded in learning science and applied to real production sessions.
  • Build simulated students grounded in real student data, and test whether they respond to good and bad tutoring the way real students do.
  • Turn expert pedagogical principles into criteria that can be evaluated reliably at scale.
  • Train models where prompting falls short, for student simulators and beyond.
  • Run empirical work end to end: form hypotheses, experiment against real data, and write up findings that change what we build.
Qualifications
  • Graduate research experience (PhD student, recent PhD, or equivalent track record) in CS, ML, NLP, or a related field.
  • Hands‑on experience with some of: LLM evaluation, post‑training/fine‑tuning, agent systems, student simulation, dataset curation.
  • Familiarity with education research (pedagogical evaluation, student modeling, tutoring dialogues, knowledge tracing, etc.) is a strong plus.
  • Clear research planning and communication, and the ability to make research legible to people outside your subfield.
You Might Be a Great Fit If You...
  • Have spent serious time teaching or tutoring.
  • Are willing to read fifty real tutoring transcripts before proposing a metric.
  • Are comfortable owning a problem where the right approach isn't known yet.
  • Treat evaluation as a research problem in its own right.
  • Would rather show a rough result this week than a polished one next month.
  • Already use coding agents heavily in your day‑to‑day work.
Why Now?

For decades, educators have dreamed of giving every student a personalized tutor. With LLMs, that vision is within reach. If you're on Team Human, there is no more meaningful or urgent mission to pursue.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Tutor Research Scientist — Evaluation & Simulation
AI Tutor Research Scientist — Evaluation & Simulation

Heyaristotle • San Francisco (CA)

Hybrid
USD 85,000 - 120,000
Research Engineer
Research Engineer

Tutorintelligence • Watertown (MA)

On-site
USD 80,000 - 120,000
Learning Scientist
Learning Scientist

Aifund • Mountain View (CA)

On-site
USD 180,000 - 260,000
Learning Scientist
Learning Scientist

career • Mountain View (CA)

On-site
USD 140,000 - 200,000
Robotics Research Scientist
Robotics Research Scientist

Tutorintelligence • Watertown (MA)

On-site
USD 80,000 - 120,000
Learning Scientist
Learning Scientist

AI Fund • Mountain View (CA)

On-site
USD 150,000 - 210,000
AI Research Scientist, Learning & Evaluation
AI Research Scientist, Learning & Evaluation

Studyfetch • Beverly Hills (CA)

On-site
USD 150,000 - 210,000
Medical, Dental, Vision (100% employer
75% dependent coverage
401(k) with employer matching
+2
Robotics Research Scientist
Robotics Research Scientist

Tutor Intelligence • City of Watertown (NY)

On-site
USD 85,000 - 110,000
AI Engineer
AI Engineer

Uncover • Mountain View (CA)

Hybrid
USD 150,000 - 190,000
Research Scientist
Research Scientist

Tessera Labs • San Jose (CA)

On-site
USD 140,000 - 210,000