REKI: Registering Knowledge in VLAs
Abstract
Most vision-language-action models (VLAs) predict actions from the current observation alone, without retaining information from the past. Existing memory mechanisms either directly pass past frames into the vision-language backbone (VLM), unroll the model during training, or add dedicated memory modules, which increase training or inference costs for VLAs. To address this, we introduce REKI, a VLA with a lightweight memory mechanism that substantially improves performance with negligible overhead. A small set of memory tokens in the action head summarises outputs from the VLM at each step and is cached across time to inform next actions. We find that memory is particularly useful for reducing repetitive robot behaviours. Evaluations on LIBERO-Plus, RoboTwin, and real-world robot tasks demonstrate REKI's strong performance and generalisation. With only 2.2B parameters and trained only on public robot data, REKI achieves state-of-the-art results among models of comparable size and can even match or surpass larger models. We release our code, data curation pipeline, and models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.