SparEdge: Redundancy-Aware Dynamic Sparse Inference with Edge-Cloud Collaboration for Multimodal LLMs
Abstract
Multimodal large language models (MLLMs) excel at cross-modal understanding and generation, yet their massive parameter size and heavy computation make inference prohibitively expensive, especially on resource-constrained edge devices. Cloud-only inference offers strong compute but suffers from high latency and bandwidth pressure. Existing acceleration schemes largely rely on static pruning or fixed sparsity patterns, and thus fail to adapt to the dynamically changing redundancy distributions in multimodal inputs. To address this, we propose SparEdge, a redundancy-aware dynamic sparse inference framework with edge-cloud collaboration for accelerating MLLM inference. First, a lightweight cross-modal redundancy-aware mechanism efficiently characterizes input-dependent redundancy across visual and textual representations with minimal additional overhead, enabling dynamic identification of redundant computation during inference. Second, an edge-cloud collaborative dynamic sparse inference mechanism translates the detected redundancy into adaptive sparsity decisions and coordinates computation between edge and cloud resources according to the runtime sparsity characteristics and resource conditions. Experiments on multiple datasets show that SparEdge improves inference throughput by more than 1.8× while maintaining accuracy, and reduces computational and memory overhead by more than 25.2%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.