Revisiting What, Where, and When in Latent Memory Generation for Frozen LLM Reasoners
Abstract
Latent memory generation offers a promising approach to augmenting the test- time reasoning of frozen large language models, in which an auxiliary Weaver synthesizes continuous memory representations and a Trigger selectively injects them into the reasoning stream. However, existing implementations rely on three convenient but weakly justified defaults: the Weaver is trained with brief ratio- nales, latent memory is read from the final Transformer layer, and memory recall is considered only at punctuation boundaries. We revisit these choices through a unified What–Where–When framework. For what to remember, we propose CoT- Supervised Memory Compression (CSMC). CSMC first trains the Weaver with richer CoT supervision and then introduces task-driven compression during Trig- ger reinforcement learning: CoT supervision improves 15 of 16 evaluated con- figurations, while compression further improves eight using fewer memory vec- tors. For where to extract, we propose Depth-Aware Readout (DAR) and show that intermediate-layer readout outperforms final-layer readout in 11 configura- tions and matches it in one, with a macro-average gain of 4.50 percentage points. For when to recall, we introduce the Reasoning-State-Aware Trigger (RSAT), which replaces punctuation-based heuristics with predictive-entropy-based acti- vation and improves StrategyQA accuracy from 63.17% to 76.13% on Qwen2.5- 7B-Instruct. Together, these results recast latent memory generation not as a fixed architecture but as an optimizable design space spanning what information is dis- tilled and retained, where memory is read out, and when it is recalled, providing a more systematic foundation for augmenting the reasoning of frozen language models. The code will be released at https://github.com/anonymous-researcher- coder/latent-memory.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.