acceptodds
Under review as a conference paper at ICLR 2027

LAnoDet: Self-Supervised Lung Anomaly Detection with Transformer-Based Diffusion Models

Abstract

Detecting anomalies in chest X-rays is central to the timely diagnosis of lung disease, but labelled abnormal images are scarce and the abnormalities are often subtle. Reconstruction-based generative models trained only on normal images offer a label-free alternative in which an anomaly is revealed by the error between an image and its reconstruction. We identify three limitations of current diffusion-based detectors: convolutional U-Net backbones capture local texture but not the long-range structure of the thorax, whereas global self-attention costs quadratic time in the number of patches; the noise-prediction objective and stochastic sampler inject sampling variance into the very residual that serves as the anomaly score; and a fixed timestep code gives the network no control over how diffusion time is represented. We propose LAnoDet, in which each component answers one of the above limitations: (1) a time-embedded Transformer encoder–decoder whose attention is factorised into row and column multi-head attention (RC-MHA), reducing the cost of attention from O(N^2d) to O(N^1.5d) while preserving structured spatial dependencies; (2) a deterministic reverse process that reconstructs the clean image at every timestep with a mean-squared-error objective instead of predicting and subtracting noise; and (3) a learnable timestep embedding added to the patch and positional embeddings. Trained only on normal chest X-rays from COVIDx CXR-4, LAnoDet reaches 96.2% accuracy, 97.3% sensitivity, 96.9% F1 score and 96.51% AUC on COVID-19 detection and transfers without retraining to a pneumonia dataset with 96.78% accuracy, 98.34% sensitivity, 97.71% F1 score and 97.16% AUC, matching or exceeding supervised classifiers reported on these datasets. A sixteen-variant ablation shows that row–column attention, skip connections and the learnable time embedding each contribute, the full model exceeding its global self-attention counterpart by 4.17 AUC points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.