acceptodds
Under review as a conference paper at ICLR 2027

Cross-Modal Patch Distillation with Budget-Enforced Any-Resolution Routing for Long-Context Compression

Abstract

Large language models (LLMs) are severely bottlenecked by long-context inputs during inference due to the quadratic complexity of self-attention and linear Key-Value cache growth. In this work, we introduce Cross-Modal Patch Distillation (CMPD), a context compression framework that maps extended textual contexts into a continuous cross-modal patch manifold prior to the transformer feedforward layers, under the hypothesis that visual manifolds can represent abstract information more efficiently than textual manifolds. CMPD leverages a model's native multimodal pathways to construct budget-constrained 2D patch grids, transforming long documents into a compact sequence of continuous cross-modal latent patches. To preserve critical verbatim tokens (e.g., proper nouns, dates, and numerical identifiers) without spatial low-pass degradation, CMPD dynamically routes query-relevant lexical anchors via a parameter-free Query-Centroid Anchor Routing mechanism. To support dynamic token budgets at inference time without retraining, we optimize CMPD in two stages: (1) multimodal manifold anchoring via a cross-modal contrastive objective on dense captioning data, and (2) causal task distillation across a diverse mixture of long-context benchmarks with stochastic budget sampling. Extensive evaluations across Qwen3.5-0.8B, Qwen3.5-9B, and Gemma-4-E4B on LongBench-v2, ZeroSCROLLS, and RULER demonstrate that CMPD achieves up to context compression, consistently outperforming discrete prompt compression and pruning baselines (such as LLMLingua) on complex multi-hop reasoning and long-form generative summarization while matching or exceeding uncompressed full-context performance. CMPD establishes a scalable and theoretically grounded paradigm for efficient long-context reasoning in modern foundation models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.