AI Research Scientist, Learning & Evaluation

Studyfetch

Beverly Hills (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, Dental, Vision (100% employer
75% dependent coverage
401(k) with employer matching
Daily team dinner in-office
Mission-driven small team

Job summary

Studyfetch Beverly Hills, CA, is seeking an AI Research Scientist focused on learning and evaluation. You will build the evaluation layer under model training and learning science across Learn Engine and the Honen platform, driving metrics that matter for student outcomes.

You’ll own multi-turn tutoring benchmarks, evaluate models in production, and align research with product decisions in a founding-team setting.

Qualifications

  • PhD in statistics, computer science, machine learning or quantitative field with 5+ years applying it to real products or research.
  • Experience evaluating AI products in production and a strong causal reasoning mindset.
  • Ability to design, run and defend experiments and benchmarks for AI models.
  • Strong communication skills and clarity in writing and speaking.

Responsibilities

  • Own the evaluation framework for the models we train and deployed.
  • Design and run evals against candidate models, including safety and reliability checks.
  • Build an internal benchmark for tutoring conversations and publish results when appropriate.
  • Bridge product data and model training to drive measurable improvements.

Skills

Evaluation for AI products
Experimental design
Causal inference
Bayesian methods
LLM evaluation

Education

PhD in statistics, CS, ML or related field

Tools

MongoDB
PostgreSQL
Vector databases
GCP

Job description

AI Research Scientist, Learning & Evaluation

Studyfetch Beverly Hills, California, United States


About this position

About Studyfetch


StudyFetch is the #1 AI-native learning platform globally, transforming how millions of students learn through personalized AI-powered education. We’re growing fast with backing from top-tier investors and a mission that’s redefining the future of education and ethical learning.



Why this role exists

We're a technology company building AI-native learning products used by more than seven million students worldwide, alongside Honen, our workforce-learning platform for organizations. Both run on the Learn Engine, the intelligence that moves a learner from initial understanding to demonstrated mastery. We work with partners like NVIDIA to bring responsible, learning-first AI to the students who need it most.


Nobody has settled how to measure whether an AI tutor teaches. Public benchmarks tell you a model can answer a question. They don't tell you whether a fourteen-year-old understood the explanation, whether the model handed over the answer when it should have asked a follow-up, or whether the student could still do the problem a week later. We train and tune our own models for learning outcomes rather than leaderboard scores, and that only works if someone can define what better means and prove when we've hit it.


That's this role. You'll build the evaluation and measurement layer that sits under model training, product decisions, and learning science across both StudyFetch and Honen. This is a founding-team role. You'll work directly with the people making the decisions, and the standard you set is the one every model and feature gets held to.



What we believe

Every learner deserves the chance to succeed. StudyFetch started with one idea: high-quality, personalized learning should be within reach for anyone, at any stage of life. Honen carries that belief into the workforce.


Accessible to everyone.


Meet people where they are.


Learning never stops.


We hire people who share this conviction. The work is demanding and the hours can be long, and what sustains you through it is caring whether a real student finally understands the material.



What you'll own

Evaluation for the models we train.


We fine-tune our own model family for tutoring, and you decide how we know whether a new checkpoint is better than the last one. That covers accuracy and reasoning, and it also covers child safety, resistance to sycophancy, and whether the model teaches Socratically instead of answering outright. You'll design the evals, run them against every candidate model, and hold the release bar.


An internal benchmark for multi-turn tutoring.


Single-turn Q&A benchmarks miss almost everything we care about. You'll build and maintain our benchmark for real tutoring conversations, decide what it measures, and defend those choices to researchers outside the company. Expect to publish parts of it.


The link between product data and model training.


The Learn Engine records what worked for past students at the personal, course, topic, and global level. You'll turn that into training signal and into evidence: which interventions moved mastery, which ones only moved engagement, and which model behaviors correlate with a student actually learning.


Data quality for expert-verified content.


We build assessment questions with subject-matter experts, starting in nursing licensure and expanding into medical and legal. You'll measure agreement between experts, catch where the official answer key is out of date, and design how those verifications feed back into training.


Analytics across both products.


Retention, activation, feature adoption, conversion, and how each of those moves when a model changes. You'll build the dashboards and reporting that product and leadership actually use, and you'll say plainly when the data can't answer the question yet.


Instrumentation we don't have.


You'll find the missing events, telemetry, and logging, then work with engineering to add them. Most of the interesting questions here are currently unanswerable because nobody logged the right thing.


The measurement bar for the team.


How we run experiments, what counts as a result, when a change ships. The patterns you set are the ones the rest of the team follows.



What we're looking for

You're a strong fit if either of these is true:



  • A PhD in statistics, computer science, machine learning, economics, physics, computational social science, or another quantitative field, plus 5+ years applying it to real products or research, or

  • Fewer credentials on paper and a track record of owning evaluation or measurement for an AI product that shipped to real users. Show us the work.


Beyond that:


You've evaluated LLMs in production, not just read about it.


You're rigorous about causality.


You write and speak clearly.


You use AI every day and have informed opinions about it.


The mission is why you're here.



The stack you’ll work in

You don't need every item below, but you should be deep in most and able to ramp quickly on the rest:



  • Statistics: experimental design, causal inference, Bayesian and frequentist methods, hypothesis testing

  • AI/LLM: eval frameworks, LLM-as-judge and its failure modes, RAG, embeddings, agent workflows, fine-tuning and post-training

  • MongoDB and PostgreSQL, vector databases, warehouse and pipeline tooling

  • Infra: GCP, and comfort reasoning about inference cost, latency, throughput, and GPU utilization

  • Reporting: dashboards people return to, in whatever tool gets there fastest


Learning science, psychometrics, or item response theory is a real plus. So is having worked with children's data and the rules that come with it.



What to expect

This is an in-person role at a fast pace, with periods of intense work around major launches.


You'll have significant ownership and autonomy with limited oversight. The role suits researchers who do their best work with room to run.


You'll be the first person in this seat. Some weeks are model evaluation, some weeks are a retention question from the founders, and you'll have to decide which one matters more that week.


It's a strong fit for people who have shipped analysis that changed a decision. If your experience has been mostly reports that nobody acted on, this likely isn't the right match.



  • 100% employer-paid Medical, Dental, and Vision; 75% dependent coverage

  • 401(k) with employer matching

  • Daily team dinner provided in-office

  • A small, mission-driven team changing how the world learns


#LI-SF1

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Research Scientist, Learning & Evaluation
AI Research Scientist, Learning & Evaluation

Socket.dev • Beverly Hills (CA)

On-site
USD 180,000 - 280,000
Daily team dinner provided in-office
Staff Engineer, Applied AI
Staff Engineer, Applied AI

StudyFetch • Beverly Hills (CA)

On-site
USD 170,000 - 270,000
100% employer-paid Medical, Dental, and Vision
75% dependent coverage
401(k) with employer matching
+1
Senior Data Scientist, Education
Senior Data Scientist, Education

Learning Commons • Redwood City (CA)

On-site
USD 190,000 - 261,800
401(k) employer match
Paid volunteer time off
Relocation support
AI Analysis Specialist
AI Analysis Specialist

Lockedinai • New York (NY)

Remote
USD 80,000 - 120,000
Equity options
Flexible remote work
Impactful work on a widely used product
Software Engineer, Full Stack
Software Engineer, Full Stack

Uncover • Mountain View (CA)

Hybrid
USD 170,000 - 210,000
Senior Software Engineer
Senior Software Engineer

MetAntz • Palo Alto (CA)

On-site
USD 200,000 - 300,000
Founding AI Engineer
Founding AI Engineer

Worky • San Francisco (CA)

On-site
USD 225,000 - 255,000
Research Engineer — Reinforcement Learning
Research Engineer — Reinforcement Learning

firecrawl • San Francisco (CA)

On-site
USD 180,000 - 270,000
Up to 0.15% equity
Generous PTO
Parental leave
+2
Founding Engineer, ML + Forward Deployment
Founding Engineer, ML + Forward Deployment

Autostep • San Francisco (CA)

On-site
USD 140,000 - 210,000
LLM credits
Startup benefits
Founder access
+3
Research Engineer
Research Engineer

Firecrawl • United States

Hybrid
USD 210,000 - 275,000
Salary that makes sense
Competitive equity
Generous PTO
+6