RILO: Feeding Implicit Process Rewards Back into Visual Reasoning for Industrial Anomaly Understanding
Abstract
Industrial anomaly understanding requires reasoning over subtle, localized, and category-dependent visual evidence whose relevance may change as the response evolves. However, standard multimodal large language models (MLLMs) typically encode the input image only once and keep visual-token representations fixed throughout decoding. This creates a mismatch between the evolving visual demands of anomaly reasoning and the static visual representation used for generation. To bridge this gap, we propose RILO (Reward-guided Inspection Loop with Online Visual Modulation), a parameter-efficient post-training framework that uses implicit process rewards not only for policy learning but also as online feedback for visual adaptation. RILO learns state-aligned token-level rewards from answer correctness, evidence compatibility, and structured-output consistency, and periodically uses them to adapt query-image representations during response generation. This generate–evaluate–revise loop redirects visual inspection toward evidence relevant to the current reasoning state, without requiring annotated reasoning traces or additional spatial supervision. Extensive experiments demonstrate that RILO achieves state-of-the-art accuracy on the MMAD benchmark under both one-shot and zero-shot settings, reaching 81.26 and 79.52, respectively. The code will be publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.