acceptodds
Under review as a conference paper at ICLR 2027

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding

Abstract

Large Vision-Language Models (LVLMs) map visual inputs into dense token sequences, increasing inference and memory costs. Elastic compression addresses this by training a single model that can run at multiple visual token budgets. However, existing approaches struggle under aggressive compression. Nested spatial pooling behaves as an imperfect low-pass filter and induces spectral aliasing that obscures fine-grained detail, while nested query resampling replaces explicit grid-aligned tokens with non-local summaries and substantially degrades spatial grounding. To resolve this representational conflict, we introduce PARCEL (**P**ool-**A**nchored **R**esampling with **C**onditioned **EL**astic Queries), a visual tokenization architecture that dynamically partitions the labor of feature extraction. PARCEL establishes spatial pool tokens as low-frequency layout anchors and conditions elastic query tokens on these anchors through Pool-Conditioned Query Resampling. This coupling encourages queries to capture complementary visual details while strengthening the pooled pathway itself, including at budgets where no query tokens are leveraged at inference. Extensive evaluations across 30+ benchmarks and two LVLMs show that PARCEL improves the performance-efficiency Pareto frontier, consistently outperforming existing matryoshka baselines across visual token budgets while preserving the "train once, deploy anywhere'" paradigm.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.