acceptodds
Under review as a conference paper at ICLR 2027

VideoTeach: Do Video and Instructional Quality Translate into Learning Transfer?

Abstract

Generative models increasingly solve complex multimodal reasoning problems and present their solutions as instructional videos. Yet evaluating video-based teaching requires tracing the full path from the generated video, through teaching, to changes in learner performance. Existing benchmarks typically examine only one part of this path. We introduce VideoTeach, a benchmark of 1,176 teaching cases across five academic domains that evaluates all three layers at the problem level. It measures video artifact quality, instructional process quality, and a simulated student’s learning outcomes on an anchor problem, a surface-transfer problem and a structural-transfer problem. Evaluation across four representative systems yields two findings. First, video and instructional quality are almost unrelated to learning gains (|r|≤0.05). Second, transfer remains weak: anchor-problem gains reach +8.3 percentage points, whereas surface- and structural-transfer gains peak at +4.5 and +1.2 points and can be negative. This transfer gap persists with stronger teacher backbones, including GPT-5.5, Gemini-3-Pro, and Claude-Opus-4.7. Additional teaching turns mainly improve anchor-problem performance without producing consistent surface- or structural-transfer gains. Overall, high-quality instructional videos do not guarantee transferable learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.