acceptodds
Under review as a conference paper at ICLR 2027

READ: Query-Relevant Evidence Retrieval for Visual Token Compression in Edge–Cloud VLM Inference

Abstract

Multimodal foundation models enable edge-cloud collaborative intelligence by transmitting visual tokens from edge devices to a cloud vision-language model for downstream reasoning. However, uploading dense token sequences incurs substantial communication overhead and cloud-side computation and memory costs. Vision-only selectors operate before transmission but cannot adapt to specific user queries. Conversely, query-aware methods typically require executing heavy language or multimodal networks on resource-constrained edge devices, incurring substantial local overhead. We present READ (Retrieving Evidence via Attention Distillation), a framework that reformulates pre-transmission token selection as an evidence retrieval task. During offline training, READ leverages prompt-to-visual attention from the frozen cloud model to supervise a compact cross-modal retriever that selects query-relevant visual tokens for accurate and efficient collaborative inference. By jointly training the retriever with prefix-ranking and score-calibration objectives, READ produces well-calibrated relevance scores that enable edge devices to identify task-critical visual tokens for efficient collaborative inference without running large models locally. Experiments demonstrate that READ achieves a competitive efficiency-accuracy trade-off across diverse architectures and transmission budgets, effectively preserving reasoning performance even under aggressive compression with minimal edge overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.