Defect Responses in Visual World Models: Long-Horizon Propagation, Gradient Interaction, and Planning Outcomes
Abstract
Visual world-model predictors are trained on local latent targets, whereas planners rank action sequences through composed predictions. Defect response is the separation between two rollouts that share future actions: one starts from the observed next latent, while the other replaces its non-action channels with the model prediction. We ask whether attenuating this response over long horizons improves planning. Defect-Response Fan (DRF) penalizes the terminal response and assigns predictor-parameter credit only through the injected prediction. A full-gradient all-suffix control (FG-AS), inspired by ACPC's same-action pairing, penalizes every suffix response and differentiates through both branches. We freeze every component except the predictor and compare both objectives with Native one-step training on PushT and PointMaze using five paired seeds, a training-inaccessible candidate bank, and matched 1,200-update and 1,800-second total-wall views. Under matched wall time, FG-AS has lower mean-regret point estimates than Native by 4.57% on PointMaze and 7.03% on PushT, despite reaching fewer updates; at 1,200 updates, the direction reverses on both tasks. All paired regret intervals include zero, so the prespecified cross-task, dual-budget confirmation gate fails. Across 40 method–Native contrasts, three response measures attenuate in 23–25 cases, but their Spearman correlations with true-regret improvement are -0.148, 0.052, and 0.032, with cluster-stratified intervals spanning zero. DRF and FG-AS retain 24.6% and 17.7% of Native optimizer throughput. In this setting, the objectives measurably alter response propagation, but response attenuation is not established as a reliable surrogate for planning quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.