acceptodds
Under review as a conference paper at ICLR 2027

LatentWeave: Rethinking Context Compression through Residual Memory Adaptation

Abstract

Long-context inference imposes growing demands on computation, memory capacity, and data movement. Soft-token compression reduces these costs, yet preserving relevant information does not guarantee its effective use by a frozen language model. We decompose excess predictive log-loss into compression-induced information loss and predictive mismatch, showing that lossy compression can improve prediction when the reduction in mismatch outweighs the information lost. Guided by this analysis, we introduce **LatentWeave**, a residual memory adaptation framework that jointly optimizes memory construction and representation adaptation. LatentWeave couples a compressor with a shared low-rank residual adapter that applies nonlinear, content-dependent corrections to the complete history-memory prefix, including embeddings from uncompressed segments. Target-derived key-value (KV) and attention supervision initialize the adapter, after which both components are jointly refined through the frozen target. On LongMemEval, LatentWeave achieves higher mean overall accuracy than the uncompressed Qwen3-8B target across all evaluated nominal assistant compression factors from to , while reducing end-to-end inference latency by 39.3–49.5%. Across Qwen2.5-14B, Qwen3-8B, Qwen3.5-4B, and Qwen3.8-27B, LatentWeave also consistently outperforms the leading soft-token compression method LatentPress in mean overall and user-fact accuracy, with relative gains of up to 27.5% in mean overall accuracy. Component ablations further show that joint compressor–adapter refinement achieves lower reference-answer log-loss than updating either component alone. LatentWeave brings memory construction and representation adaptation into a unified framework, allowing context compression to improve predictive quality while reducing inference cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.