Less Data, Smaller Batches: When Implicit Regularization Helps Informed Meta-Learning
Abstract
Informed meta-learning predicts on a new task using its few observed examples and task-specific knowledge. When the training pool contains only a few tasks, a model may learn spurious correlations between that knowledge and the observations that fail under distribution shift. Small example batches can discourage reliance on such shortcuts in supervised learning; we ask whether small task batches help when entire tasks are scarce. A finite-pool analysis shows that, for vanilla SGD, smaller task batches increase a local penalty on disagreement among task gradients. We analyze why Adam does not inherit the same scalar result, then evaluate both optimizers. With Adam, the small-versus-full-batch predictive log-likelihood benefit is much larger with 25 training tasks than with 200, on both controlled sinusoids and a prospective benchmark constructed from measured weather. This supports a scarcity-dependent optimization effect; the experiments do not identify a particular shortcut. Separately, pre-materializing tasks and overlapping data transfer yield an approximately end-to-end pipeline speedup on one GPU. Code: https://github.com/thirdperson11111/iclr_38898.gitgithub.com/thirdperson11111/iclr_38898.git.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.