AI Tutors Fail to Meet Tutoring Criteria

I love the idea of independent research checking the claims of AI tutoring! Here’s an interesting qualitative study on whether AI tutors actually do what we would call tutoring. The answer here appears to be no. Even with prompting, the tutors don’t meet the criteria of what tutoring should look like. The only promising thing here is that at least the AI models didn’t provide (many) false answers for the simulated learners.

Can AI Tutors Actually Teach? We Measured It We ran 779 simulated tutoring conversations to test whether AI models can teach, not just answer. We simulated an 8th-grade math student with a planted misconception (drawn from documented algebra errors), six frontier models acting as the tutor, and a rubric for what counts as teaching. Three things stood out: 1. Out of the box, models don't tutor. Without a tutoring prompt, they handed the student the answer in 97% of conversations. Zero successful tutoring outcomes in 298 tries. 2. A tutoring prompt helps, but unevenly. Answer-giving dropped sharply, but no model was good at both halves of the job, diagnosing the misconception and holding back the answer. 3. Price didn't predict quality. The cheapest model (about $0.002 per conversation) scored best on the strict rubric. One that costs 12x more did worse. The models were rarely factually wrong. We found 3 false statements across all 779 conversations. The problem is that handing over answers isn't teaching, and there's no standard way for buyers to tell the difference ahead of time. Funders are putting a lot of money into AI tutoring. Independent, ongoing evaluation is inexpensive by comparison, and it's the gap they're best positioned to close. Full methods, results, and caveats: https://lnkd.in/gJC9NhyV

To view or add a comment, sign in

Explore content categories