acceptodds
Under review as a conference paper at ICLR 2027

ForeSense: Learning Latent Visual Intuition for Vision-Language-Action Models via Future Experience Distillation

Abstract

The success of reasoning in large language models has inspired growing efforts to equip vision-language-action (VLA) models with similar reasoning capabilities. Existing approaches, however, typically externalize such reasoning either through verbal chains of thought or through explicit predictive representations of future evolution. We argue that reasoning for action generation should take the form of an abstract intuition about future dynamics, without being constrained to explicit representational formats or predefined semantic or temporal roles for intermediate reasoning states. To this end, we introduce ForeSense, a VLA framework that learns to sense how a scene may evolve by distilling privileged future experience into latent visual intuition. ForeSense is built around three key designs. 1) A Future-Conditioned Attention Bottleneck channels information from future observations through an autoregressive latent chain, which is then distilled into a current-only reasoner to form latent visual intuition without predefined semantic or temporal roles. 2) To support progressive refinement within this latent reasoning process, Detached Latent Curriculum (DeLaC) gradually extends the chain while applying stop-gradient to the existing latent prefix, encouraging each new step to refine previously learned reasoning states while mitigating latent-chain collapse. 3) To translate latent visual intuition into actions, Reasoning-to-Action Fusion (ReAct-Fuse) integrates the frozen reasoner’s latent key-value memory into each action-expert layer alongside the native VLA attention pathway, enabling fusion across heterogeneous reasoner and VLA backbone architectures. Extensive experiments on simulated and real environments demonstrate the strong and consistent performance of ForeSense, with success rates of 97.3% on LIBERO, 85.9% on LIBERO-Plus, and 79.2% on real-world tasks. Notably, the largest robustness gain appears under language perturbations, where success rises by 4.9 points from 86.3% to 91.2% over the π0.5-LIBERO base policy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.