acceptodds
Under review as a conference paper at ICLR 2027

Smooth Actions, Stale Targets: Temporal Weighting in Chunked Policies

Abstract

Learned robot policies can output commands that change abruptly from one step to the next, even when they predict several future actions together. Where smooth control is required, post-hoc filters aim to reduce these changes while retaining task success. However, earlier commands may no longer suit the current action: weighting the previous direction heavily can delay a turn. Each command has a target time: the time it is intended to execute. We ask whether task success depends on how filters weight commands for different target times, even when they attenuate the same input equally. We compare two filter pairs. Within each pair, the filters use the same four positive weights in opposite order, emphasizing either the newest or the oldest command. This reversal leaves their frequency-wise attenuation on a common input exactly unchanged. On one fixed OpenVLA with Optimized Fine-Tuning (OpenVLA-OFT) checkpoint, the recent-weighted filters achieve success rates and percentage points higher than their past-weighted counterparts. Each comparison uses paired resets, matched by task and initial state. A second checkpoint on a different task suite shows the same pattern. The resulting closed-loop jerk and dynamics differ within each pair. Equal attenuation therefore need not imply equal task success. Standard filters illustrate the practical implications of this distinction. Across two Octo streams, a causal four-sample moving average over the latest selected commands (SW(4)), which returns a linearly changing command's value from control steps earlier, lowers pooled success by percentage points. In contrast, a four-sample causal Savitzky–Golay endpoint filter (SG4), which preserves the current value on linear command streams, reduces jerk while retaining the unsmoothed aggregate success. Tuned exponential moving averages offer useful trade-offs on OpenVLA-OFT; selected adaptive One Euro settings likewise exhibit a smoothing–success trade-off. When post-hoc smoothing is required, consider both attenuation and how much the filter emphasizes earlier commands, and judge the trade-off by closed-loop success.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.