Modulation-Only Feature Reconstruction for Unified Anomaly Detection
Abstract
Feature reconstruction for anomaly detection uses the difference between the input features and a reconstruction learned from normal data. We propose Modulation-Only Feature Reconstruction (MOFR), which pairs an image-independent initial state with per-token feature conditioning. The decoder of MOFR starts from learnable positional embeddings shared by all images, and the input features generate the modulation coefficients of the residual updates and of the output head instead of forming the initial state. Using the per-token conditions and a global condition obtained from their mean, the decoder reconstructs features without token-to-token attention. On MVTec-AD, VisA, and Real-IAD, MOFR reaches image-level AUROCs of 99.8%, 99.0%, and 91.1% and AUPROs of 96.1%, 95.8%, and 96.3%. Crossing the initialization pathway with per-token conditioning shows that the contribution of per-token conditioning depends strongly on whether the initial state contains the input features. The differences from feature-stream reconstruction, which uses an input-derived state, are small and vary with the dataset and the metric. These results show that pairing a shared initial state with per-token conditioning is a reconstruction scheme that replaces input-derived initialization with conditional updates while retaining strong anomaly detection performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.