Foundation Model Alignment Regularizes Transformer-Based Unrolled MRI Reconstruction
Abstract
Deep unrolled reconstruction networks alternate data consistency steps with learned refinement networks and are typically supervised only at the final reconstruction. More expressive refinement networks have produced limited improvements over the convolutional E2E-VarNet on fastMRI, partly, we hypothesize, because their training converges to suboptimal solutions. Across single-coil knee and multi-coil brain reconstruction, we find that the bottlenecks of unregularized transformer cascades collapse early to a small fraction of their available rank and do not recover, whereas the convolutional model partially recovers during training. We address this problem by aligning the refinement network's bottleneck of every cascade with dense representations extracted from the fully sampled ground truth by a frozen diffusion or self-supervised foundation model. A learned convolution in each cascade maps the bottleneck to the corresponding foundation model features. The foundation model and the alignment convolutions are removed after training, leaving the deployed reconstruction network and its inference cost unchanged. We evaluate this approach on single- and multi-coil datasets using multiple foundation models and convolutional and transformer-based reconstruction architectures. At longer training budgets, the aligned transformers surpass aligned E2E-VarNet, whereas their unregularized counterparts remain below unregularized E2E-VarNet. On knee data, the gains increase with acceleration and are largest for the transformer architectures. Alignment maintains high-rank bottlenecks across cascades and improves transformer training when unaided optimization stalls or deteriorates. These gains arise from foundation model alignment rather than increased rank alone: directly regularizing the bottleneck rank raises it beyond that of every aligned model without producing a lasting improvement in reconstruction. Where tested, the improvement over unregularized training is significant at in one-sided paired Wilcoxon signed-rank tests over evaluation slices, with Holm correction across the seven reported tests.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.