Improving Certified Few-Shot Adaptation without Changing Clean Predictions
Abstract
Few-shot predictors can be misled by correctly labeled copies of their support examples. Copying changes class counts and can reverse predictions even when class-conditional feature distributions stay fixed. We separate count-dependent coefficients from support statistics and hold these coefficients at their initial values. This normalization preserves every clean prediction for trimmed-distance and exponential-cache scores, with certificates for their complete modified scores. For the cache, we prove that normalization cannot reduce robustness to a single cross-class stored-source copy when all sources are mutable. The remaining copy sensitivity depends on within-class kernel deviations. Across eight text and image encoders, normalization reduces copy harm without changing any clean prediction. On two image tasks, normalized Tip-Adapter retains 87.5% clean accuracy, and 85.4% of predictions remain correct under every single cross-class stored-source copy, compared with 80.4% for zero-shot classification. Under the broader threat of four arbitrary bounded-feature replacements, certified accuracy rises from zero to 28.3%. Fixing initial coefficients can thus preserve adaptation gains while improving resistance to support-set poisoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.