S²-Memory: A Zero-Training Sparse Concept Key-Value Memory
Abstract
All AI and cognitive systems rest on a shared and rarely questioned assumption: memory requires storage and training. From Transformer KV caches to the classic "memory trace" theory in cognitive science, this assumption has shaped our understanding of memory. We have identified a simple quantity that governs the retrieval behavior of sparse concept representations. We call it the truncation loss Z: the difference in retrieval quality between a concept activation that is fully retained and one truncated to its top-k entries. For the system we study, Z is large and decisive—retaining the full activation raises MRR from 0.47 to 0.91 (Z = 0.44)—which identifies truncation, not the choice of representation, as the dominant failure mode. To examine this, we built S²-Memory—a 126MB, zero-training, plug-and-play system that represents each sentence as a sparse key over a frozen semantic center space. It retrieves over 156,660 sentences from 54 novels of widely varying genre and style using half the storage of an fp32 dense index (120.3 MB vs. 240.6 MB). Integrated with the 1.5B Qwen model (Qwen Team, 2024), it answers cross-book questions without knowing which book contains the answer. Training cost: 1,000. Examining the system revealed three further results. First, the self-referential state that motivated its design—a decayed accumulation of concept activations used directly as the retrieval query—is not merely unhelpful but harmful, and should be removed (MRR 0.0014 against a random baseline of 0.0021). Second, the concept centers are highly collinear (effective rank 292.7 of 2000; mean pairwise cosine 0.199); ZCA whitening repairs this and raises R@5 from 0.830 to 0.920. Third, the abstract centers can be replaced by content-word centers that preserve retrieval quality (R@5 0.925 vs. 0.930) while making every center nameable: 99.6% of activated centers are semantically related to the sentence, and masking the content-word centers collapses retrieval quality (MRR 0.791 → 0.234) whereas masking function-word centers costs almost nothing (MRR 0.791 → 0.788). On our corpus, BM25 is strongest under lexically overlapping queries (MRR 0.9942) but degrades sharply under non-overlapping queries (MRR 0.2650, median rank 1 → 21), where dense retrieval leads (MRR 0.7240). Our sparse key does not exceed dense retrieval in either condition. Because the intermediate representation is a set of natural-language concept labels rather than a dense vector, it can enter a language model's context directly, without a trained probe or projection layer, and it is not bound to any particular encoder; we do not evaluate the downstream benefit of this interface.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.