TutorMoments: Do AI Tutors Know When to Help and When to Hold Back?
Hugging Face
Hugging Face introduces TutorMoments, a framework to evaluate whether LLMs can balance helping students versus letting them struggle. Built on real tutoring data, it uses teacher-annotated key moments to test AI tutors in simulated sessions. Preliminary results show models over-help by default, but explicit prompting improves performance, though still lagging behind human judgment.
Hugging Face has released a preview of TutorMoments, a framework designed to measure whether cutting-edge LLMs can balance the trade-off between helping a student and holding back to encourage independent reasoning. The framework is built on a dataset of 462 de-identified transcripts from real one-on-one math tutoring sessions with U.S. students in grades 2-7, annotated by experienced math teachers who flagged key moments where tutors had to choose between scaffolding and pushing for rigor. In tests, seven LLMs were given transcripts up to a decision point and asked to continue as the tutor with a simulated student. Results showed that with a plain prompt, models tended to over-help, but when the trade-off was explicitly described in the prompt, all models improved, though still falling short of human tutors' consistency. The researchers are releasing the dataset, code, and replays for reproducibility, and emphasize that automated evaluation cannot replace real learning outcome studies. The project received support from the Gates Foundation and Learning Commons.
- Abbreviations
- LLM = Large Language Model — большая языковая модель
Source: Hugging Face blog —
original
