Organize, Then Read: Rethinking Token Mixing via Structured State Representation
Abstract
Self attention has become a central token-mixing operator in embodied policies, yet many alternatives primarily address the quadratic cost of pairwise interaction rather than how source information should be organized before retrieval. We introduce Structured State Representation (SSR), an operator based on the principle of organize first, then read. Within each head, learned state centers shared across samples serve as common addresses for source organization. Source tokens are competitively assigned and aggregated into input-dependent states, after which queries selectively retrieve from the resulting compact state bank. This separation yields linear token-mixing complexity for a fixed number of states without constructing dense query–source interaction matrices. Controlled operator-replacement experiments in ACT and Diffusion Policy show that SSR remains competitive with standard attention and improves performance in selected settings, including human demonstrations. Qualitative latent analyses further suggest more progress-consistent representation geometry while preserving high correspondence with Transformer representations. Together, these findings support structured state organization as a viable alternative to direct pairwise attention in embodied policy learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.