Gains Are Gains? What drives accuracy gains in transductive CLIP methods?
Abstract
Transductive CLIP methods leverage unlabeled test batches to improve recognition, with performance gains typically summarized by aggregate classification accuracy. Yet the same accuracy gain can arise for very different reasons: a method may reduce errors caused by classes that are absent from the batch, improve discrimination among the classes that are present, or do both. We introduce a decomposition that separates these two effects and complement it with controlled interventions that vary the unlabeled context while holding key aspects of the prediction problem fixed. Across multiple fine-grained recognition datasets and CLIP variants, the resulting analysis reveals substantial differences that aggregate accuracy obscures. Methods with similar final performance can behave very differently. Moreover, a considerable fraction of the gain can remain even when the adapted scoring function is replaced by frozen CLIP scores while preserving the class set induced by the method. The controlled interventions further show that the balance between these effects depends strongly on the composition of the unlabeled context. Together, these findings show that final accuracy alone incompletely characterizes transductive adaptation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.