acceptodds
Under review as a conference paper at ICLR 2027

-0.6: A Unified Vision-Language World Model with Adaptive Streams-of-Thought

Abstract

Intelligence in the physical world requires self-directed, task-adaptive reasoning through different representations of the world as a problem unfolds. We introduce Ψ-0.6, a unified vision-language world model that adaptively selects the representations through which it reasons. Ψ-0.6 reasons through streams-of-thought: sequences interleaving images and text with geometrically-grounded, interpretable visual abstractions such as camera pose, depth, optical flow, and point tracks. A shared token interface supports any-to-any prompting, allowing the same model to generate, interpret, and compose these representations as intermediate reasoning steps. Ψ-0.6 matches state-of-the-art zero-shot performance on MindCube, a challenging benchmark of multi-view spatial reasoning. Composing these capabilities — estimating camera motion, selecting the relevant observed view, and answering from it — further improves performance. World modeling enables counterfactual imagination: the model translates a specified object motion into a flow prediction, generates the resulting scene, and answers from it. On FlowBench-Real, which tests this ability, Ψ-0.6 shows better generalization than Qwen. Through reinforcement learning from task feedback, Ψ-0.6 learns adaptive streams-of-thought that select visual representations for each input and task. This improves visual question-answering accuracy, with gains transferring to unseen benchmarks. By unifying perception and imagination, geometry and language, as operations of the same model, Ψ-0.6 uses them for representation-adaptive reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.