acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Readouts in CLIP-Based Zero-Shot Anomaly Detection

Abstract

CLIP-based zero-shot anomaly detectors build task-specific decision rules on top of frozen vision-language representations. However, a strong representation does not guarantee a strong decision: useful anomaly evidence may already be encoded in the representation but not fully used by the readout, i.e., the mechanism that turns the representation into an anomaly score. We study this gap in CLIP-based zero-shot anomaly detection. For a canonical binary CLIP readout, the anomaly score depends on a single decision direction. We show that substantial predictive information remains outside this direction, that the remaining information is structured rather than a by-product of high dimensionality, and that the useful directions vary across anomaly contexts. We then ask whether this unused information can be recovered with only a simple change to the detector. We keep the original detector frozen and add a lightweight, image-conditioned reader that makes only a small bounded correction to the existing decision pathway, using auxiliary anomaly data for training. Despite its limited capacity, this intervention improves a diverse set of CLIP-based anomaly detectors. Our results suggest that improving the representation is only part of the problem in zero-shot anomaly detection. It is also important to ask how much task-relevant information the final readout actually uses, and whether information left unused by the original decision can be recovered.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.