acceptodds
Under review as a conference paper at ICLR 2027

LoRA Readout: Learning to Read Parametric Memories for Context-Free Inference

Abstract

Encoding context into model parameters can reduce the inference costs of in-context learning, but conventional fine-tuning can be costly and require context-specific supervision. We introduce LoRA Readout, a framework that separates context writing from memory reading to enable context-free inference. The writer encodes raw context into a memory LoRA through next-token prediction, without requiring context-specific question-answer pairs or other auxiliary supervision. A shared reader is trained across frozen memories to access the encoded knowledge and use it to answer queries. Incorporating a new context requires only an offline memory-writing step, while the reader is reused without further adaptation. At inference, the model combines the memory and reader to answer queries without including the original context tokens in the prompt. Experiments across multiple datasets show that LoRA Readout outperforms related approaches with a small training-token budget: only M training input tokens for source-memory writing and reader training on Qwen3-8B, less than of the smallest context-training corpus reported by these approaches.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.