acceptodds
Under review as a conference paper at ICLR 2027

Mass-Preserving Visual–Text Fusion for Training-Free Few-Shot Class-Incremental Learning

Abstract

Training-free Few-shot Class-incremental Learning (FSCIL) must use scarce support images without confusing class-specific evidence with probability reallocation between base and novel classes. This distinction matters because a globally normalized interface can change both the ranking within the novel group and the competition between novel and base classes at the same time. Direct support-image fusion therefore makes the source of an endpoint change difficult to identify. We introduce a Mass-preserving Visual–text Fusion (MP-VTF) correction at the Bilevel Modality Calibration (BiMC) interface to separate these effects. MP-VTF uses labeled support images during frozen inference, without updating model parameters or accessing query labels. It applies a query-wise log-sum-exp translation that retains the fused within-novel ordering while restoring the text branch’s original novel-group probability mass. Across CIFAR100, CUB-200-2011, and miniImageNet, Naive fusion increased the mean novel-group probability mass in all 78 session–seed records, whereas MP-VTF restored it to numerical precision. MP-VTF produced small positive novel-class changes, nearly unchanged overall accuracy, and a small base-class cost. The miniImageNet endpoint direction is reported descriptively because its visual weight was selected with label-dependent metrics. These results establish mass preservation as a testable interface-control principle for separating within-group evidence updates from group-level probability allocation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.