acceptodds
Under review as a conference paper at ICLR 2027

Same Objective, Different Moments: Residual-Guided Temporal Resampling for VLA Continuation

Abstract

Continued adaptation of task-adapted vision-language-action (VLA) policies seeks to improve performance under a finite update budget without inducing overfitting. Temporal resampling can reallocate this budget across trajectory moments, but without correction it also reweights the baseline continuation objective, making it difficult to determine whether performance gains arise from temporal allocation or objective reweighting. We introduce Residual-Guided Temporal Resampling (ReTime), which decouples temporal exposure from objective weighting. ReTime uses Flow-Matching residuals from a frozen starting policy to construct a non-uniform temporal proposal and applies importance correction to preserve the baseline continuation objective and its expected gradient. It modifies only the temporal sampling distribution, while keeping source and trajectory selection probabilities fixed. Across Florence and Qwen backbones in simulation and real-world settings, ReTime consistently achieves the strongest average performance. On Florence–SimplerEnv, 500 additional updates improve performance by 22.92 percentage points over the starting policy and by 11.20 points over the strongest continuation baseline. Over longer-horizon real-world continuation, ReTime continues to improve after Uniform begins to degrade. Mechanistic analysis further shows that, under matched temporal allocation, importance correction provides an additional 4.69-point gain on Florence–SimplerEnv, demonstrating that temporal exposure and objective weighting are not equivalent and motivating their explicit decoupling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.