acceptodds
Under review as a conference paper at ICLR 2027

Can Large Language Model Agents Learn by Watching Tutorial Videos?

Abstract

Humans continuously acquire knowledge from carefully crafted materials such as books and tutorial videos, carrying that understanding into future tasks. Millions of such instructional videos are uploaded online every day, yet LLM agents remain unable to exploit this ever-expanding wealth of human teaching. Their capabilities are frozen after training, while parameter updates at scale remain prohibitively expensive. This raises the central question: can LLM agents learn from this vast instructional resource without parameter updates? We introduce Tutorial Agentic Learning (TAL), a nonparametric continual-learning paradigm that consolidates tutorial videos into a persistent, structured tutorialbook memory through two-stage adversarial consolidation. At inference time, TAL retrieves through a Table-of-Contents (TOC) tree, tying retrieval directly to the model's reasoning rather than to embedding similarity. To study this setting, we introduce VidLearn, a multimodal QA benchmark built from 72 chess tutorial videos spanning 12 openings. On VidLearn, TAL achieves up to 113% and 91% relative gains on the training and test splits, improves as more videos are consumed, and yields stronger opening positions in 69% of head-to-head games. It also outperforms the evaluated retrieval and memory baselines, while a separate software-tutorial evaluation demonstrates cross-domain generalization. Together, these results highlight instructional videos as a scalable resource for parameter-update-free continual agent learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.