Freeing Activation Memory for Task Adaptation of Pathology Foundation Models: Layer-Selective Freezing and Partial Low-Rank Adaptation
Abstract
Task Adaptation of Pathology Foundation Models (TAPFM, kumar2025tapfm) showed that a vision-transformer (ViT) pathology foundation model (PFM) and an attention-based multiple instance learning (MIL) aggregator can be jointly fine-tuned on a single GPU by detaching gradients between two separate computational graphs. This makes end-to-end adaptation practical without the multi-GPU clusters used by prior work, but the number of tiles that can be sampled per whole-slide image (WSI) per training step remains small (e.g., 100 tiles for GigaPath), because activation memory for backpropagation through a giant ViT scales with both the number of tiles and the number of transformer blocks. Motivated by the observation in TAPFM that task-adaptation gradients concentrate in the final transformer blocks with negligible updates to earlier layers, we translate this empirical finding into an actionable memory-saving mechanism. We introduce and compare three single-GPU strategies for GigaPath: (i) low-rank adaptation (LoRA) applied to all transformer blocks, which reduces gradient and optimizer-state memory; (ii) freezing all but the last transformer blocks, which additionally removes the frozen prefix from the autograd graph and eliminates its activation memory; and (iii) a hybrid that freezes all but the last blocks and applies LoRA only within those unfrozen blocks, combining both memory savings. Empirically, these strategies raise the maximum tiles-per-WSI-per-epoch achievable on a single H200 from TAPFM's original 100 to 300 (LoRA-only), 500 (freeze-last-N only), and 1000 (freeze-last-N + LoRA) with , respectively. We provide a memory-decomposition analysis that explains why LoRA alone does not reduce activation memory while layer freezing does, formalize the hybrid strategy as a generalization of TAPFM's detached dual-gradient update restricted to a block subset, and show across institutional and external cohorts that the proposed configurations retain strong FGFR3 and EGFR mutation-prediction performance while offering task-dependent trade-offs between predictive accuracy and resource efficiency. Our results suggest that removing activation-memory bottlenecks is a viable path toward more scalable and parameter-efficient single-GPU adaptation of large PFMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.