Compress at Input, Recover at Depth: Improving the Efficiency–Performance Trade-off for MLLMs
Abstract
In common MLLM pipelines, full-token connectors (e.g., Linear or MLP projectors) forward all modality tokens into the decoder, making attention cost grow quadratically with input length. Token compression alleviates this cost but can weaken multimodal conditioning, creating a core efficiency–performance trade-off. This issue is particularly pronounced for long or information-sparse scientific modalities, where aggressive compression can discard task-relevant input information. We propose ENCORE, a dual-stage connector built on a simple principle: compress at input, recover at depth. ENCORE constructs a compact set of instruction-conditioned query tokens to reduce the number of input modality tokens, and introduces a lightweight one-shot recovery route that uses intermediate text states to retrieve complementary information from pre-compression modality memory. A stabilized residual update integrates this information during prefill, with no repeated memory reads during cached decoding. Across vision, protein, and molecule tasks, ENCORE achieves the best overall Norm-Avg against the shared MLP and Q-Former references while improving the performance–efficiency trade-off. Relative to MLP, it reduces inference FLOPs by 16.3–62.6% and lowers training time in four of five settings. Extended comparisons further show that ENCORE remains competitive with stronger compression and multi-layer reinjection baselines while retaining a compact decoder interface. Ablations show that compression provides a strong primary representation, while one-shot recovery contributes complementary gains, especially under aggressive compression. Overall, ENCORE provides a practical, modality-agnostic connector that consistently improves the performance–efficiency trade-off across diverse settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.