A Video Question Answering Dataset for Validating Understanding of Scientific Goal-Directed Cognition
Abstract
In scientific research, understanding experiments in the laboratory calls for not only physical scene understanding, but also for inferring how actions relate to a scientist's goals, assessments, and evolving plans. We introduce LabThought, a dataset of naturalistic egocentric video from in-situ wet lab research in organic chemistry, biochemistry, and microfluidics. LabThought also includes concurrent think-aloud narration during experiments and post-experiment interviews, providing unusually direct evidence about the reasoning and intentions underlying behavior in the laboratory. We annotate these transcripts to highlight statements related to different components of goal-directed cognition; specifically assessments of the experimental state, rationales behind decisions, and intentions. Using these annotated transcripts, we construct a video question-answering evaluation in which questions are grounded in the expressed cognitive states. Initial experiments with multimodal language models show substantially different performance across question types and context conditions: concrete intentions are generally easier than questions about rationales and assessments, while providing explicit action descriptions improves performance. These results suggest that LabThought provides a useful setting for studying how well video models connect observed laboratory activity to the underlying goal-directed cognition.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.