InfiniteWorld: Separating Visual Evidence from Familiar Physics
Abstract
Physical reasoning requires adapting familiar expectations to the dynamics of the observed world. Existing benchmarks establish broad physical capabilities, yet an overall score leaves unclear whether answers follow visual evidence when familiar rules no longer apply. We introduce InfiniteWorld, a benchmark that pairs physical-reasoning questions with simulator videos whose physical parameters are systematically altered. Its generation pipeline connects simulation, scene screening, paired video translation, question construction, and input checks. Five tasks assess law identification, parameter estimation, temporal reasoning, analytic counterfactual comparison, and law induction. The evaluation defines five matched input settings for each question, comprising text-only evaluation and ordered or shuffled observations from simulator and translated videos. We evaluate a diverse set of video-language models under these matched settings. The results expose a persistent preference for familiar physical rules. Frames improve performance for most models, but gains differ across tasks. Strong accuracy on Earth-normal questions coexists with poor recognition of altered dynamics, and Earth-normal answers remain common when the observed physics changes. Task scores and semantic agreement further distinguish matching physical factors from matching their amount, and correct answers from repeated errors. InfiniteWorld makes these behaviors measurable through paired changes in the evidence supplied for the same question, separating visual evidence from familiar physics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.