Learning from Consequences: Posterior-Transition Reweighting for Behavior Cloning
Abstract
Robot demonstrations contain more than action labels: they also record the observations that follow each action. We investigate whether these future observations can guide how action supervision is allocated during vision-language-action (VLA) post-training. Posterior-Transition Reweighting (PTR) identifies the observed future of each demonstrated action chunk among alternatives drawn from the training data. The same identification model provides an auxiliary representation-learning objective and a bounded, stop-gradient weight on the original action loss. This separates two uses of future information: learning features and weighting supervision. PTR requires neither reward labels nor an evaluable action likelihood and applies directly to flow-matching regression. Our analysis shows that the expected score equals the identification information carried by the candidate list minus the scorer's error, and that the bounded weights redistribute supervision without discarding any example. Experiments on LIBERO, RoboCasa, and twelve real-robot tasks examine heterogeneous pooled training and controlled corruption of demonstrations. In pooled training, PTR raises real-robot success from 50.0% to 63.8% over standard fine-tuning, with gains on all twelve tasks. Controls in pooled simulation training show that reweighting contributes beyond memory and auxiliary representation learning. These results support future observations as a practical source of additional supervision for VLA adaptation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.