acceptodds
Under review as a conference paper at ICLR 2027

Verified Supervision and Compositional Generalization in a Finite Program Space

Abstract

This work studies whether low-rank fine-tuning on exhaustively enumerated, automatically verified composition trajectories can change a small language model's probability distribution so that correct programs rank higher on new tasks under unchanged evaluation. The experiments use Qwen3-0.6B for polynomial transformations over finite fields, where a solution is a sequence of three atomic operations chosen from five. All 125 distinct programs were enumerated and scored exactly, so Hit@32 uses no sampling or repeated candidates. With matched optimiser-step and loss-bearing target-token budgets, mean Hit@32 rose from 45.32% for atomic control to 72.72% for composition fine-tuning, a gain of 27.40 percentage points (95% interval [25.40, 29.40], ), positive in all six paired continuations from one atomic checkpoint. This contrast measures the full training recipe. Prespecified secondary comparisons showed a gain of 26.13 points on familiar compositions in new fields and losses of 16.02 and 11.80 points on the two withheld-pair families. Atomic APPLY accuracy decreased in all six composition branches relative to their paired controls. Selective pair exclusion also lowered withheld-stratum Hit@32 across 13 completed subsets. Experiments on Qwen3-0.6B-Base found no clear advantage of structured exploration over independent sampling at comparable candidate diversity, and 400 steps of training with a verifiable reward changed held-out coverage by +0.22 points (95% bootstrap interval [, ]). A development-set follow-up with three continuation seeds from one checkpoint showed the same transfer pattern with higher mean program entropy. The positive effect supports consolidation of familiar composition families and transfer to new tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.