Towards Faster Anomaly Detection and Reasoning via KV Cache Compression
Abstract
Industrial anomaly detection (IAD) has become a crucial technology for identifying anomalies in complex real-world environments. Recently, multimodal large language models (MLLMs) have demonstrated outstanding global visual understanding and zero-shot cross-domain reasoning capabilities across general multimodal tasks. As a result, several studies have begun applying MLLMs to IAD-related tasks by integrating domain knowledge and specialized methods to enable causal reasoning about anomalies, achieving competitive results on relevant benchmarks. However, these approaches often overlook the high computational cost associated with MLLM inference and the substantial key-value (KV) cache redundancy introduced by high-resolution images. This study first investigates the attention distribution patterns of MLLMs in the IAD domain. Our observations reveal that models predominantly focus their attention on normal regions, while only a few KV pairs corresponding to anomalous areas receive effective attention—potentially leading to incomplete perception of anomalies. To address this, we propose AnomalyKV, a novel KV cache compression method that combines anomaly awareness with efficient inference. Our approach adaptively allocates KV cache budgets at a fine-grained level, enabling instance-level dynamic compression based on differences among input samples. Extensive experiments show that AnomalyKV minimizes performance degradation across multiple mainstream MLLMs while significantly improving inference efficiency and reducing inference time costs. Notably, AnomalyKV achieves 2.57× real-time acceleration in full-scenario tasks, matching the performance of models with full KV caches and substantially outperforming existing mainstream KV cache compression methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.