acceptodds
Under review as a conference paper at ICLR 2027

Teaching MLLMs to Think Like Quality Inspector via Hierarchical Reasoning

Abstract

Driven by the rapid evolution of foundation models, the paradigm of industrial anomaly detection is shifting from traditional deep learning approaches to Multimodal Large Language Models (MLLMs). While MLLM-based methods have demonstrated impressive capabilities, existing works often fail to satisfy the requirements of real-world industrial production. Current research predominantly focuses on determining merely whether a defect exists, while neglecting that industrial quality inspection and subsequent process adjustments rely heavily on diagnostic details, i.e., which object is involved, why a defect is suspected, and what judgment is made. Inspired by the diagnostic process of human inspectors, we abstract object and defect morphologies into generalized structured knowledge and propose Hierarchical Knowledge Preference Optimization (HKPO). Through hierarchical branch-level and segment-level preference optimization, the model learns to identify objects, reference structured knowledge, and judge defects, thereby ensuring a rational and well-grounded quality inspection process. Extensive experiments on public benchmarks show that our framework lifts a lightweight 8B model 7.15% above its own backbone and 2.71% above a 4 larger 32B counterpart, surpasses the strong proprietary model by 5.46% in average accuracy, and outperforms dedicated anomaly-detection MLLMs by at least 7.95% on anomaly discrimination.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.