Meta-Weighted Self-Training for Language Models under Extreme Data Scarcity
Abstract
Adapting language models from only a small set of prompt–completion pairs is challenging: direct fine-tuning overfits, while self-training on model-generated data can reinforce errors. We propose , a meta-weighted self-training framework for extreme data scarcity that assigns each synthetic example a weight reflecting its : examples are upweighted if imitating them leads, after several adaptation steps, to the largest likelihood improvement on the real demonstrations. The weights therefore assign credit over future training dynamics rather than ranking examples by current sample quality. Starting from a frozen pretrained backbone with LoRA adapters, MWeiST generates synthetic completions and optimises these weights via a meta-objective on the held-out real demonstrations. To make this tractable at scale, we introduce , a first-order approximation that replaces backpropagation through the inner optimisation with gradient alignment between per-example updates and the meta-loss; a double-gradient implementation recovers all alignments in two backward passes per step, independent of batch size. Working from as few as 80 real demonstrations, we adapt five backbones spanning 270M to 4B parameters on three summarisation benchmarks. MWeiST attains the lowest held-out perplexity in every model–dataset cell and improves completion cross-entropy over all five baselines in the large majority of paired comparisons. GradAlign tracks the exact BPTT meta-gradient closely while admitting substantially larger inner batches at the same memory budget. Learning which synthetic examples to train on is therefore practical at scales where exact meta-gradients are not.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.