acceptodds
Under review as a conference paper at ICLR 2027

The Missing Temporal Dimension in Flow-Based Action Chunking

Abstract

Action chunking has emerged as a cornerstone technique in robotic control and imitation learning. Recent robotic manipulation policies predominantly employ action-chunked policies trained with flow matching objectives, achieving substantial performance gains. Despite this success, existing approaches represent action chunks merely as unstructured vectors during flow-based generation, overlooking the physical-time dependencies among actions within each chunk and potentially leading to ineffective policy learning. In this paper, we show that additionally supervising the first-order temporal derivatives of actions within a chunk provides a surprisingly effective way to incorporate physical-time dependencies into flow-based policies. We further establish that this supervision scheme corresponds to a special case of path-independent flow matching, in which action chunk generation evolves jointly along two dimensions: denoising time and physical time. By capturing the joint effects of multi-step denoising and the physical-time evolution of actions, our method enables fully exploiting the rich temporal structure of action chunks while requiring only one additional flow-based loss term. Across a broad range of simulation benchmarks and real-robot experiments, our method consistently improves both performance and sample efficiency over the base policy, with particularly pronounced gains on real-robot tasks. These results highlight the importance of explicitly modeling the temporal structure inherent in action chunking.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.