acceptodds
Under review as a conference paper at ICLR 2027

SparEdge: Redundancy-Aware Dynamic Sparse Inference with Edge-Cloud Collaboration for Multimodal LLMs

Abstract

Multimodal large language models (MLLMs) excel at cross-modal understanding and generation, yet their massive parameter size and heavy computation make inference prohibitively expensive, especially on resource-constrained edge devices. Cloud-only inference offers strong compute but suffers from high latency and bandwidth pressure. Existing acceleration schemes largely rely on static pruning or fixed sparsity patterns, and thus fail to adapt to the dynamically changing redundancy distributions in multimodal inputs. To address this, we propose SparEdge, a redundancy-aware dynamic sparse inference framework with edge-cloud collaboration for accelerating MLLM inference. First, a lightweight cross-modal redundancy-aware mechanism efficiently characterizes input-dependent redundancy across visual and textual representations with minimal additional overhead, enabling dynamic identification of redundant computation during inference. Second, an edge-cloud collaborative dynamic sparse inference mechanism translates the detected redundancy into adaptive sparsity decisions and coordinates computation between edge and cloud resources according to the runtime sparsity characteristics and resource conditions. Experiments on multiple datasets show that SparEdge improves inference throughput by more than 1.8× while maintaining accuracy, and reduces computational and memory overhead by more than 25.2%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.