acceptodds
Under review as a conference paper at ICLR 2027

DexForesight: Self-Distilled Foresight for Dexterous Vision-Language-Action Models

Abstract

Task-oriented dexterous manipulation requires coordinated arm–hand actions whose quality is determined by the physical and visual consequences that follow them. Demonstration trajectories therefore contain more supervision than the demonstrated actions alone: they also record how objects move, articulate, and appear after each interaction. We introduce DexForesight, a training framework built on a single observation: a demonstrated future can supervise a causal policy in two fundamentally different ways—a predicted future teaches the representation what interaction state to anticipate, while the realized future teaches the action generator how its current prediction should change. For representation learning, we adapt V-JEPA 2 to demonstration trajectories and design an action-conditioned latent predictor with structured arm–hand controls; its predicted future latent supervises learnable query tokens in the vision-language model (VLM) representation. For action generation, we execute the same policy under causal and privileged future context; their velocity difference defines a future-conditioned flow correction learned by a lightweight residual adapter. Future observations and simulator object states are used only during training and are removed at deployment. Across the 11 single-arm and bimanual tasks of DexJoCo, DexForesight improves the average success rate of the DexJoCo π0.5 baseline by 10.1 percentage points (62.6% vs. 52.5%). Both supervision paths improve the policy in isolation, and their combination performs best; combining physical and visual privileged future context also achieves the strongest average performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.