acceptodds
Under review as a conference paper at ICLR 2027

Lexical Conditional Memory: Decoupling Specialized Memory from Reasoning for Knowledge-Intensive Tasks

Abstract

LLMs couple specialized lexical memory and general reasoning within a shared parametric backbone, creating two challenges for knowledge-intensive tasks. First, domain-specific terminology is often fragmented into multiple native subtokens, potentially requiring repeated composition across layers before higher-level reasoning. Second, the specialized vocabularies required by knowledge-intensive tasks are often large, long-tailed, and continuously evolving. Encoding and maintaining such lexical knowledge entirely within shared model parameters is therefore inefficient. To address these challenges, we introduce Lexical Conditional Memory (LCM), a retrofit architecture that externalizes specialized lexical memory while preserving the native tokenizer and pretrained Transformer backbone. LCM comprises two complementary mechanisms. Deterministic variable-length lexical routing recognizes complete domain terms across variable-length subtoken spans and maps them directly to external memory entries, providing direct access to complete-term lexical representations. Hierarchical conditional memory injection integrates the retrieved representations through depth-specific interfaces, enabling specialized lexical knowledge to be accessed on demand rather than stored uniformly in the shared parameters. Together, these mechanisms provide an explicit external pathway for specialized lexical representations, reducing reliance on their implicit reconstruction and storage within the shared backbone. Across six benchmarks in medicine, law, and chemistry, LCM outperforms Engram by 2.1–2.2 macro-average points across two controlled Qwen2.5 scales and by 2.4 points in the native-Engram Argon setting. LCM also remains complementary to factual retrieval, with LCM+RAG achieving the highest overall macro-average performance. These results establish external lexical memory as an effective approach for adapting LLMs to knowledge-intensive tasks across architectures and domains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.