TopAct: Topology-Aware Action Tokens for Autoregressive Robot Policies
Abstract
An action tokenizer determines both the information retained about a robot trajectory and the categorical prediction problem presented to its policy. We introduce TopAct, an ordered codec that represents a 16-step action chunk with four finite-scalar-quantized tokens. Beyond reconstruction, its training objective aligns neighborhoods in action space with neighborhoods in latent space, before and after quantization. We evaluate matched topology-aware and reconstruction-only (Recon) codecs through action-only codec training, compact autoregressive policies, and an OpenVLA integration, holding the codec frozen while training each policy. TopAct preserves substantially stronger neighborhood structure after quantization and produces policy targets that are consistently easier to predict: token accuracy is higher for the compact policy on every suite, and under OpenVLA the TopAct arm converges faster and to a stronger token model. The same targets also support better closed-loop behavior: the compact policy reaches higher success on three of four suites, and the selected OpenVLA policy attains a higher success point estimate. Together, the findings motivate topology-aware action representations and the joint evaluation of codec quality, policy learning, and closed-loop behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.