acceptodds
Under review as a conference paper at ICLR 2027

Act on the Low Frequencies: Prefix Execution of Vision-Language-Action Chunks

Abstract

Vision-language-action policies predict chunks of robot controls and ship every chunk in full, and how much of a chunk a robot needs to finish its task has remained unmeasured in closed loop. We introduce frequency-prefix execution (FPE), a training-free scheme that transforms each finished chunk into a temporal cosine basis, transmits a fixed number of its lowest-frequency rows, and lets the robot reconstruct and execute the whole chunk under the unchanged policy, horizon and replanning schedule. The basis makes the discarded detail exactly accountable: the reconstruction error of a prefix equals the energy of the omitted rows, and under a contraction condition the accumulated state deviation stays within the largest single-chunk residual. On the official LIBERO protocol with OpenVLA-OFT, two cosine rows keep 98.8% success on Spatial against 98.6% for quantized full codes at 54% of the native payload, and four rows keep 92.6% against 92.2% on LIBERO-10 at 72% of the payload, both passing a registered non-inferiority test. On Goal the tolerance to truncation follows the task, which motivates PairGap, a shortest-path certificate that bounds, at a single restored state, the regret of executing a prefix against the best option in a fixed bank of prefixes and full codes. Together they show that a robot can act on a fraction of each predicted chunk, and they identify the tasks and states that call for the full code.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.