acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Zero-shot Evaluation in Continual Vision-Language Pretraining

Abstract

Continual pretraining of vision-language models is commonly evaluated by zero-shot accuracy, which informs whether to retrain from scratch every year at about four times the compute or to keep all past data for replay. Accuracy drops during training are read as forgetting. Yet zero-shot accuracy conflates the image features with their alignment to class-name embeddings, while also depending on evaluator-chosen class names and templates. We disentangle representation quality from zero-shot alignment on all 28 released TiC-CLIP checkpoints by comparing zero-shot classification with nearest class mean and linear probing on the same frozen image features, while varying only the prompt. The checkpoints cover three training strategies: retraining from scratch on all data (Oracle), replaying all past data while continuing from the previous checkpoint (Cumulative), and sequential updating without replay (Sequential), across 2016–2022 and two data-filtering regimes. We find that zero-shot comparisons between training strategies are unstable, imprecise, and inflated, and that zero-shot declines alone do not establish representation forgetting. With model weights fixed, switching between two commonly used templates reverses the sign of the Oracle–Cumulative gap on 6 of 11 datasets under basic filtering. Across datasets, paired tests fail to distinguish Oracle from Cumulative using zero-shot accuracy under either data filter, but separate them using linear probes under both. Under the probe, replay is worth 0.6 to 1.5 points, a third to a fifth of the 3.2 to 4.4 that zero-shot accuracy reports, and its value grows at most a third as fast over time; retraining from scratch adds a further 0.6 to 0.7 points. Along the Cumulative trajectory, where the training pool only grows, zero-shot accuracy drops by at least one point on 10 of 132 one-year steps, and the linear probe on none of them. These findings motivate three checks for comparing continual-pretraining strategies from frozen checkpoints: probe the image features alongside the zero-shot head, test paired per-dataset differences rather than suite-level averages, and verify conclusions across standard prompt configurations. Without these checks, differences attributed to training may instead reflect the evaluation protocol.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.