SwiftAD: Visibility-Guided Hierarchical Token Pruning for Efficient Zero-Shot 3D Anomaly Detection
Abstract
Zero-shot 3D anomaly detection must localize subtle defects without target-domain training data while remaining efficient enough for practical inspection. Existing CLIP-based methods render point clouds into multiple high-resolution views and process all visual tokens, even though many tokens correspond to empty backgrounds, regions without valid 3D support, or redundant normal surfaces. We present SwiftAD, an efficient framework centered on visibility-guided hierarchical token pruning. SwiftAD first converts rendering-derived 3D-to-2D correspondences into a geometric visibility mask and removes unsupported background tokens without additional learnable parameters. It then applies a lightweight anomaly-aware selector to retain informative tokens from the remaining object regions. We find that token pruning introduces a distributional shift in the visual features, weakening their alignment with the original CLIP text prototypes. Textual prototype adaptation therefore learns separate global and local prototypes for object-level recognition and point-level localization, respectively, and applies residual calibration to their anomaly prototypes. Experiments on MVTec3D-AD, Eyecandies, and Real3D-AD show a favorable accuracy–efficiency trade-off. SwiftAD preserves strong zero-shot performance while halving computational cost and improving throughput by 73.9% over the unpruned baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.