DualFuture: Event-Triggered Dual-Future Residual Correction for Vision-Language-Action Policies
Abstract
Action-chunked vision-language-action (VLA) policies execute a predicted chunk before observing its consequences, so a small prediction error at a critical action event goes uncorrected, accumulates over the executed commands, and fails the episode. We introduce Dual-Future Residual Correction, an event-triggered framework that wraps a pretrained VLA without fine-tuning its weights. At a detected event, it predicts two futures—one consistent with task success, one induced by the proposed action—and decodes their discrepancy into a bounded residual added to the action before execution. Both futures and the residual are learned from realized rollout futures and known action differences, so no manual correction labels are needed. Instantiating critical events as persistent gripper transitions, we attach DualFuture to the baseline VLA policy on RoboCasa365 and report three further frozen policies to situate its base success. Success rises from 73.8% to 78.7% on atomic tasks and from 45.4% to 50.9% on unseen composite tasks. On five tabletop tasks with a real Franka robot, DualFuture raises the frozen policy's pooled success from 40.0% to 52.0% under randomized object placements.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.