Evidence before deployment

Before we trust
AI tutors, measure
whether they can teach.

TutorBench is an open benchmark for observable tutoring behavior. We examine how models diagnose, guide, adapt, and support learning in specified scenarios.

Open source Research-oriented Developer Preview
CASE 03 / 48Upper Elementary · Mathematics

Adding Fractions

I think 1/3 + 1/4 = 2/7. Can you help me check it?
Authored learning objective

Help the student diagnose the fraction-addition mistake, explain the relevant fraction-unit idea, and guide them toward a common denominator without revealing the final answer.

Open the full case Illustrative case walkthrough

Five dimensions of tutoring

More than
right or wrong.

We examine observable tutoring behavior using structured rubrics and transparent evaluation. Each dimension captures a distinct aspect of a response in a specified benchmark scenario.

Explore the evaluation method

Diagnosis

Whether it identifies the learner’s actual error, gap, or reasoning issue.

Dimension 1 / 5

Look beyond
the final answer.

Measurement infrastructure for AI tutoring

Open data.
Transparent evaluation.
Observable behavior.

Synthetic cases
48
Public development scenarios
Authored rubrics
128
Case-specific evaluation criteria
Current dataset
0.2a.6
tutor-eval-v0.2a
Evaluator version
0.3a.4
Open and reproducible
Developer Preview · No calibrated public model runs yet. Explore the community
Evidence & limitations

Public model results are unavailable. Human calibration (P5) has not started, and Judge-vs-human and statistical validation are not completed. Calibration infrastructure exists, but real Community Review and human calibration have not started. Judge-vs-human validation and statistical validation are not completed. TutorBench measures observable tutoring behavior in specified benchmark scenarios, not long-term learning, retention, transfer, satisfaction, or classroom outcomes.