acceptodds
Under review as a conference paper at ICLR 2027

Key-Gram: Extensible World Knowledge for Embodied Manipulation

Abstract

Language-conditioned manipulation requires both reusable task knowledge and state-dependent visual reasoning. Most vision-language-action policies represent both within a shared backbone or conditioning path, making targeted knowledge extension difficult. We introduce Key-Gram, a conditional-memory framework that extracts short, task-specific key-grams, maps them to learnable memory entries by deterministic multi-head hashing, and injects the retrieved entries into selected Transformer layers through token-wise gating and lightweight convolutional fusion. The memory is organized as independently addressable sub-tables, allowing capacity to grow without changing the backbone architecture and supporting constant-time lookup with respect to total table size. Across RoboTwin2.0, LIBERO/LIBERO-Plus, and real-world dual-arm manipulation, Key-Gram improves both and . Relative gains are on the hard RoboTwin2.0 split, on LIBERO-Plus transfer without target-domain fine-tuning, and on real-world long-horizon tasks. In four-stage continual compositional assembly, Stage 4 terminal-node success rises from to . These results support external linguistic memory as a useful architectural bias for compositional grounding and transfer under task expansion.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.