acceptodds
Under review as a conference paper at ICLR 2027

Recognition by Imagination: A Generative Approach to Form-from-Motion

Abstract

We study the problem of recognizing actions from a minimalist motion stimuli with no appearance cues. We frame this as an instantiation of Johansson’s famous experiments in 1973 on recognizing actions from a video showing a handful of moving keypoints against a black background. While this collection of keypoints seems to be spread arbitrarily in a single frame, their motion enables action perception. We establish that modern MLLMs are not able to represent such signals. The objective of this paper is then to enable video models to be sensitive to such subtle motion stimuli. To perceive this minimalist motion and recognize actions, we propose a novel two-step Recognition by Imagination (RBI) approach: (i) imagination: we fine-tune a video generation model to generate a realistic video solely from the keypoint stimuli, (ii) recognition: then, we employ a standard MLLM to classify the action. RBI significantly outperforms frontier MLLMs and shows superior generalization compared to supervised fine-tuning of open MLLMs. Interestingly, while RBI is fine-tuned on keypoint videos, it exhibits some surprising emergent capabilities. It works off-the-shelf with novel stimuli such as silhouettes, shape outlines, and even event-based camera inputs. Finally, we also show that the same model, fine-tuned only on keypoint videos, qualitatively generalizes to other psychophysical form-from-motion stimuli.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.