acceptodds
Under review as a conference paper at ICLR 2027

Meta-Weighted Self-Training for Language Models under Extreme Data Scarcity

Abstract

Adapting language models from only a small set of prompt–completion pairs is challenging: direct fine-tuning overfits, while self-training on model-generated data can reinforce errors. We propose , a meta-weighted self-training framework for extreme data scarcity that assigns each synthetic example a weight reflecting its : examples are upweighted if imitating them leads, after several adaptation steps, to the largest likelihood improvement on the real demonstrations. The weights therefore assign credit over future training dynamics rather than ranking examples by current sample quality. Starting from a frozen pretrained backbone with LoRA adapters, MWeiST generates synthetic completions and optimises these weights via a meta-objective on the held-out real demonstrations. To make this tractable at scale, we introduce , a first-order approximation that replaces backpropagation through the inner optimisation with gradient alignment between per-example updates and the meta-loss; a double-gradient implementation recovers all alignments in two backward passes per step, independent of batch size. Working from as few as 80 real demonstrations, we adapt five backbones spanning 270M to 4B parameters on three summarisation benchmarks. MWeiST attains the lowest held-out perplexity in every model–dataset cell and improves completion cross-entropy over all five baselines in the large majority of paired comparisons. GradAlign tracks the exact BPTT meta-gradient closely while admitting substantially larger inner batches at the same memory budget. Learning which synthetic examples to train on is therefore practical at scales where exact meta-gradients are not.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.