OCP-R: Stabilizing Option-Level Readout for VLM Anomaly Detection
Abstract
VLM anomaly-detection benchmarks such as MMAD and AnomalyCoT test in- dustrial inspection with multiple-choice questions, matching the closed defect vo- cabularies of industrial quality assurance, so a frozen VLM must turn its hidden state into answer-option scores, a step we call the option-level readout. In this paper, we show this step is unreliable: frozen VLMs often hold anomaly evidence that their option logits do not expose. Two measurements fix the repair. Weak frozen readouts spend most of their MCQ supervision on answer-format rather than content tokens, so what is missing is content, not capacity; and a replace- ment readout discards decisions the frozen head already scores correctly, so the correction must add to the frozen logits rather than replace them. We propose Option-Conditional Residual Readout (OCP-R), a small trained head that scores each candidate option against the frozen state and adds the result to the frozen logits, zero at initialization so training departs from the frozen readout only where supervision demands. Across nine general-purpose and four anomaly-specialized VLMs, OCP-R raises mean MMAD accuracy from 62.17% to 78.27% while pre- serving 93% of the frozen model’s correct predictions, and on AnomalyCoT it is the only method with no catastrophic held-out drop on any backbone, against one for internal-weight adaptation and three for a replacement readout.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.