Exploring the Potential of MLLMs for Zero-Shot Anomaly Segmentation
Abstract
Existing multimodal large language model (MLLM)-based methods for zero-shot anomaly segmentation (ZSAS) rely on dedicated segmentation components, leaving open whether a single MLLM can directly segment and describe anomalies. To investigate this question, we empirically examine two potential bottlenecks. First, at the vision encoder level, pretrained visual features exhibit weaker semantic relationships within regions than those of visual foundation models, potentially hindering anomaly localization that requires semantically coherent representations. Second, at the LLM level, contextual [SEG] representations exhibit limited diversity in their variation across input images under repetitive mask supervision, potentially restricting adaptation to different anomaly characteristics. Based on these findings, we propose MIAS (MLLM Itself for Anomaly Segmentation), an MLLM-only framework with two training strategies that address the respective bottlenecks. To strengthen semantic relationships in the vision encoder's features, Patch Affinity Distillation (PAD) transfers pairwise patch affinities from a frozen visual foundation teacher. To encourage more diverse image-conditioned [SEG] representations in the LLM, Complementary Prompt Supervision (CPS) constructs complementary localization tasks from existing anomaly masks without additional manual annotations. Both strategies operate only during training. At inference, MIAS uses a single MLLM to generate anomaly descriptions and predict masks by directly matching contextual segmentation queries with visual features, requiring neither external visual models nor dedicated mask decoders. Experiments across eight industrial and medical datasets show that MIAS achieves the best average segmentation performance among the compared methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.