acceptodds
Under review as a conference paper at ICLR 2027

STORC: Spatio-Temporal Output-Space Residual Correction for Robot Video Prediction

Abstract

Robot video prediction models often lose accuracy in new scenes and tasks. As target frames are gradually revealed, they expose prediction errors and provide feedback for correcting future predictions, yet reliably and efficiently estimating future errors from limited history remains challenging. To address this challenge, we propose STORC, a spatio-temporal output-space residual correction framework requiring only one forward pass of the frozen backbone per prediction clip and no online gradient updates. STORC analytically estimates local relationships between prediction changes and residual changes at a regional scale, then extrapolates future residuals along the cached prediction trajectory, anchored at the latest observed residual. To control extrapolation risk, Causal Predictive Shrinkage (CPS) evaluates propagation against residual persistence on the latest completed transition held out from fitting and adaptively shrinks the propagation increment. To preserve fine spatial detail, STORC fuses regional and native-resolution corrections using the same shrunk propagation coefficients. Under a unified causal feedback protocol, STORC achieves an average MSE reduction of 34.8% relative to the strongest evaluated adaptation baseline in each of 45 backbone-split combinations, spanning five backbones and nine test splits on BridgeData, RoboNet, and BAIR. Core latency measurements show a geometric mean speedup of 3.91 over the evaluated adaptation baselines. Small-scale offline evaluation on self-collected laboratory videos demonstrates transfer without additional training or tuning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.