Past2Next: Past Action Conditioning Policy with Data-Augmented Tokenization
Abstract
Discrete-token autoregressive policies have emerged as a compelling paradigm for robot learning. Their downstream performance depends on two axes: the action tokenizer's reconstruction fidelity and the policy's reliability in predicting action tokens. Prior work mostly targets the former, yet reconstruction gains often fail to reach downstream success, a compression gap we trace to the latter, under-addressed axis. We introduce Past2Next, a discrete-token autoregressive policy that addresses this bottleneck by using past action conditioning for token prediction, then decoding the action tokens into a continuous action chunk via a geometry-aware tokenizer. During tokenizer training, we apply data augmentation to improve codebook coverage on the \(SO(3)\) action manifold. During policy learning, Past2Next conditions token prediction on observations, a short history of executed actions, and their first- and second-order finite differences to acquire dynamic features. Our method exploits locally smooth action trajectories governed by the equations of motion. Specifically, finite-difference signals capture the local tangent and higher-order motion trends of the action stream, narrowing the feasible set of next-step actions to a compact neighborhood around its extrapolated trajectory and reducing next-token prediction uncertainty. Across simulation and real-robot experiments, Past2Next reduces predictive Shannon entropy and improves task success rates over the evaluated baselines. Moreover, by such a mechanism, increased codebook capacity improves downstream performance rather than degrading autoregressive learning, allowing finer action quantization to translate into better control performance. Modeling local dynamics and respecting action geometry thus makes an autoregressive policy both accurate and predictable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.