MOPDA: Mixed-Trajectory On-Policy Distillation for Language-Guided Industrial Anomaly Detection
Abstract
Large vision-language models (LVLMs) have recently shown strong potential for industrial anomaly detection (IAD) by providing image-level anomaly judgments and interpretable defect reasoning. However, reliably translating generated language judgments into precise pixel-level anomaly localization remains challenging. To address this, we propose Mixed-Trajectory On-Policy Distillation for Language-Guided Industrial Anomaly Detection (\MODPA), the first framework to introduce on-policy self-distillation into LVLM-based IAD. Specifically, \method introduces Mixed-Trajectory Supervision, which combines student-generated on-policy trajectories with teacher-generated evidence-conditioned trajectories under a shared token-level distillation objective for fine-grained judgment learning. To connect the learned judgment with dense anomaly perception, we further introduce Language-guided Visual Anchoring. The final judgment is used as a compact semantic condition to construct image-specific normal and abnormal anchors, which are contrasted with dense visual features to produce the anomaly map. In this way, language guides localization without directly determining the pixel-level response. Experiments on five IAD benchmarks show that \method outperforms the evaluated LVLM-based baselines on most detection, localization, and judgment metrics, while remaining competitive with CLIP-based methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.