Trust-Aware Prototype-Guided Quantum Representation Learning for Backdoor-Resistant Medical Image Classification
Abstract
Training-time backdoors turn small image artifacts into prediction shortcuts, and compact medical classifiers are especially exposed because a low-dimensional representation leaves little room to absorb a spurious direction. We present TAPQ-Net, a hybrid quantum–classical defense that trains directly on poisoned data, without a trusted clean subset, sample removal, or post-hoc purification. A frozen encoder and a fixed PCA projection map each image to eight coordinates. A label-free expectation–maximization (EM) mixture assigns every sample a soft prototype and a trust score; an 8-qubit tree-structured quantum encoder (TQE) with seven trainable rotations is aligned to that prototype, and the trust score bounds the sample's cross-entropy weight. The mechanism that separates TAPQ-Net from prior prototype- and clustering-based defenses is a reverse channel: the EM state may evolve only when the silhouette geometry of the quantum embedding stays separated, balanced, and non-collapsed. Because EM never reads labels, relabeling reaches the prototypes only through the trigger's latent offset, and alignment bounds the trigger-direction component of the embedding. On OASIS and six MedMNISTv2 datasets under random-class and targeted patch poisoning at 1–10%, TAPQ-Net attains the lowest attack success rate in all 42 conditions (0.13–0.24, against at least 0.34 for every compared method) and the highest clean accuracy in 40, and paired tests over the 42 conditions are significant for all four metrics. A topology- and parameter-matched classical tree trails the TQE by 3.4 accuracy points, poisoned samples receive lower trust without being removed, and the defense carries over to clean-label and warping-based triggers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.