From Error Rates to Educational Measures: Benchmarking Speech AI for Teacher–Child Interaction in Preschool Classrooms
Abstract
The quality of teacher–child interaction is a central determinant of early development, yet assessing it still depends on trained experts observing classrooms in person, which cannot scale. Speech AI could automate this: transcribe classroom talk, attribute it to the teacher or children, and compute the measures that interaction assessment relies on, such as talk time, turn-taking, and vocabulary. Yet speech systems are evaluated by transcription and diarization error rates, not by whether the measures they produce match those coded by humans. In this paper, we introduce MPCS, to our knowledge the first benchmark that pairs verbatim transcripts with continuous speaker-role timelines over real preschool classroom audio: **227** hours of teacher and child speech (ages 3–5) from **130** classrooms. Evaluating 14 speech recognizers and four speaker-role models, we find that children produce 16% of the characters but 39% of the strongest recognizer's errors, a gap that model scale and in-domain fine-tuning do not close. Lower error rates also do not guarantee better measures: a speaker-role model with 1.4% diarization error still overcounts teacher–child turns by 17%. What matters is the kind of error. Classroom-adapted systems measure how much teachers and children talk to within 1%, but miscount turns and overestimate children's vocabulary. Code is available at [this anonymous repository](https://anonymous.4open.science/r/MPCS-Bench-68B2).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.