acceptodds
Under review as a conference paper at ICLR 2027

VOCABSCOPE: READING CONCEPTS FROM VIDEO GENERATION STATES IN SPACE AND TIME

Abstract

Video generation models can produce rich and dynamic videos from natural-language prompts, yet their hidden states remain difficult to interpret in terms of concepts. We introduce VocabScope, a framework that maps hidden states of video generators into vocabulary space while preserving their correspondence to video locations. Our method first establishes token-level correspondence between the video generator and a frozen vision–language model, then aligns their representation spaces, and finally connects intermediate representations in its language backbone to vocabulary predictions through layer-specific Jacobian transport. This makes it possible to examine what concept evidence is readable from a generation state and associate it with locations in video space and time. VocabScope offers a direct vocabulary-space view of video-generation hidden states and provides a tool for inspecting concept evidence at corresponding video locations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.