acceptodds
Under review as a conference paper at ICLR 2027

Temporal Ensembling Revisited: Recovering the Stochastic Headroom of Generative Robot Policies

Abstract

Diffusion and flow-matching policies formulate robot control as a generative process, decoding each action chunk from sampled noise. Under closed-loop deployment, a failed rollout may therefore reflect an unfavorable sequence of noise draws rather than an inherent limitation of the policy, yet standard success rates leave the two indistinguishable. We examine this ambiguity with an episode-level pass@ probe that holds the task initialization fixed and varies only the policy noise. On LIBERO-Long, success rises from 80.7% with a single draw to 96.1% within eight, indicating that a substantial portion of long-horizon failure is sampling-induced. We recover this stochastic headroom online with Recency-weighted Temporal Ensembling (RTE), which exploits the overlapping predictions inherent to receding-horizon control, averaging them under weights that decay with prediction age. We show analytically that this weighting trades stale-plan bias against prediction noise: when predictions scatter around a plan that drifts with age, neither uniform averaging nor freshest-only execution minimizes per-step action error, and the error-minimizing weights decay with age. With a single fixed decay rate, RTE improves four flow-matching and diffusion backbones in most settings across LIBERO, RoboTwin 2.0, SIMPLER, and a real Franka arm. Controls that replan at matched strides confirm that temporal aggregation accounts for most of these gains, which replanning frequency alone cannot match, and that its contribution tracks the effective ensemble size.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.