Rethinking Batching in Visual Test-Time Learning
Abstract
In visual test-time learning, batching can change what a model learns from an image. An unrelated batch mate selects extra padding, and these tokens enter the target's inner training set even though the images never share features. Same-image interventions isolate this effect from numerical noise and ordinary padding effects. Removing all padding is not the right correction: it changes the checkpoint's singleton computation and lowers average precision (AP) by 0.592 points on a preregistered Small pilot. Canonical-Support Inner Training (CSIT) instead retains the canvas each image would occupy alone and excludes only batch-excess support from the inner loss. It requires no retraining or learned parameters and preserves the steady-state singleton output tensor-exactly. On all 5,000 Common Objects in Context (COCO) validation images, random-batch matched-object drift falls by 68–76% across two released detector scales and batch sizes 4–8. For reference scores near the decision threshold, flip rates fall from 0.30–0.32 to 0.08–0.11. When a batch may be split, exact-shape grouping is more faithful and faster in our measured workload; when a heterogeneous forward must be retained, CSIT repairs its support. AP changes are small and mixed. These findings establish inner training support as part of an explicit execution contract for inference-time learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.