KPE-AC: Kinetic Path Energy for Adaptive Chunking in Vision-Language-Action Models
Abstract
Action chunking enables a vision-language-action (VLA) policy to predict multiple future actions in a single inference, improving efficiency while maintaining temporal coherence. A main challenge is selecting the execution horizon—that is, how many actions from the generated chunk to execute before querying the policy again. Short horizons allow frequent observation updates and more precise control but require more computation, whereas long horizons reduce inference frequency but delay responses to changes in the environment. Recent works have proposed adaptive execution-horizon selection methods, but this problem is still underexplored. In this work, we look inside the generation process of a single action chunk to guide adaptive action chunking. Inspired by Kinetic Path Energy analysis in image generation, we introduce Kinetic Path Energy-guided Adaptive Chunking (KPE-AC). We adapt KPE to action generation by measuring deviations from constant-velocity transport, yielding correction KPE. We analyze how this energy reflects uncertainty in future actions and leverage it to determine how many actions to execute. Beyond looking inside a single action chunk, we assess Cross-Chunk Consistency (CCC) by comparing actions predicted for the same future time steps in the previous and current action chunks. CCC adjusts the execution horizon while leaving the predicted actions unchanged. Experiments in simulation and real-world manipulation demonstrate the effectiveness of KPE-AC across multiple flow-based VLA backbones.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.