acceptodds
Under review as a conference paper at ICLR 2027

Rulers, Not Meters: Discovering Falsifiable World Models through Active Visual Experimentation

Abstract

A useful world model must predict the effects of new interventions, so it has to be identified, not merely fit. Existing evaluations of vision-language agents fix what the agent observes and cannot tell identification from fitting guesses to feedback. We study active world-model identification in Eyes of Galileo: from raw renders of 35 counterfactual physics scenes, an agent must build its own measuring references, propose an executable symbolic world model, and revise it when execution at dispersed targets exposes counterexamples. With GPT-5.6, such an agent identifies 63.4% of hidden world models (strict or relaxed), versus at most 44.6% for three visual-reasoning scaffolds under identical feedback, and dispersed counterexamples raise exact identification from 33% to 73% in a controlled revision study. The ability depends on the backbone (33.1% with Qwen3-VL-32B, 24.0% with Qwen3-VL-8B), and so does learning it: fine-tuning on successful GPT-5.6 episodes raises the 32B model's success on unseen scenes from 42.9% to 57.1%, but does not help the 8B model. Identifying world models from pixels thus hinges on both making evidence and learning to use it, a capability that fixed-observation evaluations cannot reveal.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.