CoSTA-Diff: Medical Concept-Conditioned Diffusion for Targeted Skin Lesion Augmentation
Abstract
Diffusion-based augmentation can expand medical training sets, but class labels alone do not specify which within-class patterns should be generated. We present CoSTA-Diff, a latent diffusion framework that uses shared medical concept coordinates to connect generation conditions with augmentation targets. A frozen MONET model provides image-specific evidence for 22 dermoscopic concepts and fixed text directions. An additive conditioner combines these quantities with the diagnosis to construct class and concept tokens for cross-attention, without requiring manual per-image concept annotations. The same evidence coordinates define within-class neighborhoods, allowing sparsely represented training profiles to serve as synthesis anchors. New images are generated from fresh noise under the anchors' observed labels and concept evidence, supported by a compact wavelet latent autoencoder. In a retrospective evaluation on a retained ISIC 2018 split, adding 1,800 synthetic images yields mean ResNet-50 balanced accuracy of 71.48% and macro-F1 of 73.73% across three classifier seeds, exceeding real-only training by 2.70 and 2.33 percentage points and a class-conditioned latent diffusion baseline by 3.04 and 3.65 points. Design studies and cross-backbone comparisons characterize the conditioning choices and architecture dependence of these gains. CoSTA-Diff provides a shared representation for concept-conditioned synthesis and targeted data expansion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.