Pay for the Reasoning Once: Compiling Agent-Written Programs for Visual Anomaly Detection
Abstract
Agentic anomaly detectors invoke a vision-language model on every image they inspect, using tools that a human author must write in advance. We ask whether normal examples can instead support the construction and verification of reusable inspection code. ANP (Anchored Normality Programs) lets an agent compose visual primitives into executable checks, tests the candidates on held-out normal examples and controlled violations, and freezes the accepted checks with their fusion rule and calibration. On MVTec LOCO, the complementary ensemble ANP-E2 achieves 96.00% mean image-level AUROC, 8.50 percentage points above AnomalyMoE's reported result. The efficient ANP configuration achieves 94.34% mean AUROC. On the breakfast category, it requires 191.13 ms per image under our shared timing protocol with precomputed masks, corresponding to 8.78× and 47.32× speedups over our LogSAD and UniVAD executions. Neither configuration invokes a language model during inference. Controlled ablations of the efficient configuration attribute a gain of 2.87 AUROC percentage points to the verified program layer (95% CI [1.93, 3.84]); additional component ablations examine its role in the ensemble. Experiments across six benchmarks assess the scope of normal-only compilation, while acceptance and calibration audits characterise the evidence available before labelled defects are observed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.