TeleScopeMix: Depth-Aware Spatial and Channel Adaptation of Vision Transformers for Medical Imaging
Abstract
Adapting pretrained Vision Transformers to medical imaging is challenging when target datasets are small, imbalanced, and drawn from modalities that differ substantially from natural images. We introduce TeleScopeMix, a parameter-efficient adapter that combines depthwise spatial filtering, efficient cross-channel mixing, and a depth-dependent capacity schedule. Our analysis of supervised and self-supervised backbones indicates that adapter utilization generally increases toward later Transformer blocks, motivating greater channel-mixing capacity at depth. Across various medical-image classification datasets and Vision Transformer backbones, TeleScopeMix with optimizing fewer than of the backbone parameters achieves comparable or even better performance compared to full fine-tuning method.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.