Focus on Likely Classes for Test-Time Prediction
Abstract
We ask: Can focusing on likely classes of a single, in-domain sample improve accuracy? Prior work argued “no”, we answer “yes” on average. Standard entropy minimization yields largest gains, and we further dissect why by looking at its two core mechanisms: increasing likely and decreasing unlikely class predictions. Simply maximizing logits of the two most likely classes explains most of its benefits and, in some cases, even outperforms entropy minimization. Thus, decreasing classes is a secondary concern. Our controlled experiment and theory investigate the most puzzling case, why simply optimizing the logits of the two most likely classes leads to accuracy gains. Our small scale networks suggest that the degree of uncertainty of a prediction, followed by model choice, and to a lesser extent dataset choice moderate the success chances of the optimization. We show that the true class tends to have larger gradients independent of whether the prediction is correct or not. Our theory thoroughly investigates one of multiple possible explanations from the angle of shared features. Our evaluation focuses on 26 large language models and 12 text datasets, and 18 image recognition models trained on ImageNet. It demonstrates gains of 0.6% on average for entropy minimization on uncertain samples using no extra data aside from the given test sample using a fixed learning rate. We also suggest gradient-ray ensembling. When traversing input/embedding space by applying the same gradient with differing learning rates and aggregating probabilities for the same sample further improves accuracy significantly to 0.75% relying on a coarse range for learning rates, thus improving accuracy and eliminating the need to find a precise learning rate at the expense of computation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.