What Pretraining and Midtraining Make Learnable from Rewards
Abstract
A reward can identify a correct answer while leaving the computation needed for new inputs undetermined. We study what pretraining and midtraining contribute to this gap. In sequential state computation and contextual memory, we characterize mechanisms that agree on every training reward yet disagree on held-out answers. Task-independent source observations resolve this ambiguity. We then construct finite sampled-Adam paths from random initialization through source prediction and reward adaptation in the same parameters: prediction acquires execution or retrieval, and rewards learn how to use it. Experiments with pretrained Qwen2.5 checkpoints measure this division of labor through source midtraining. With first-operation source supervision, Sequential models acquire execution and show a strong reward-task-selection advantage over a private-random source control. Memory experiments show retrieval retention and task selection under an independent source intervention. Public tasks separate accuracy at reward entry, subsequent gain and final performance. Together, these results explain how source information becomes computation that rewards can use.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.