EdgeEdit: An Edge-Oriented Multimodal Knowledge Editing Framework for On-Device MLLMs
Abstract
On-device multimodal large language models (MLLMs) are increasingly enabling multimodal applications, yet updating their static encoded knowledge efficiently remains a critical challenge. In-context editing (ICE) provides an efficient and lightweight approach for fine-grained knowledge updates. However, on-device deployment typically demands joint model and input compression, introducing severe obstacles for ICE: model quantization impairs the model’s ability to utilize contextualized edited knowledge, while edit-agnostic token reduction inadvertently discards edit-critical visual evidence. To address these challenges, we propose EdgeEdit, an edge-oriented multimodal knowledge editing framework that coordinates offline model deployment with online inference. Specifically, EdgeEdit introduces Edit-Aware Mixed-Precision Quantization for offline model deployment, which conducts zeroth-order sensitivity profiling and budget-constrained dynamic programming to preserve contextualized edited knowledge utilization. For online inference, EdgeEdit employs EvidenceSaliency Guided Token Reduction, combining cross-modal saliency attribution with semantic relevance to progressively remove redundant visual tokens while retaining edit-critical visual evidence. Extensive experiments demonstrate that EdgeEdit maintains competitive editing performance while substantially reducing system costs. Compared with the strongest competing baseline for each efficiency metric, EdgeEdit achieves average reduction factors of 2.9×/1.5×/2.5× in memory usage, editing latency, and energy consumption on the server, and 2.0×/2.2×/2.0× on mobile devices, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.