SCOUT: CRITIC-GUIDED SCHRÖDINGER TRANSPORT FOR ONLINE VLA REFINEMENT
Abstract
Contact-rich and dynamic manipulation often requires well-crafted refinement methods that can learn from sparse rewards while preserving efficient execution beyond a large pretrained policy. We introduce SCOUT, a lightweight framework for refining pre-trained vision-language-action (VLA) policies online through critic-guided Schrödinger transport. SCOUT formulates policy improvement as a Schrödinger bridge from the VLA's reference action chunk to a distribution of promising corrections, where a critic proposes and reweights multiple candidate action chunks, thereby constructing a training target that captures coordinated changes across action space and time. We exploit the point-source structure induced by conditioning on the VLA reference action, enabling tractable bridge regression toward an online critic-weighted target. This enables simulation-free learning through endpoint regression on Brownian interpolants without iterative coupling training. We evaluate SCOUT on simulated benchmarks and real-world precise manipulation tasks. SCOUT improves success rates over RLT and DSRL baseline methods. It achieves comparable measured inference latency to RLT and improves real-world active-execution throughput over RLT and both DSRL variants. Together, these results demonstrate that reference-conditioned action transport strengthens effective online refinement with low deployment overhead. More details and visualizations can be found on our website: https://scout-vla.github.io.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.