Search Into Weights: Turning Test-Time Search into Capability per Parameter
Abstract
Test-time search produces both candidate solutions and feedback that can train a small model during inference. We propose a memory-conditioned draft model that serves as a speculative proposer and a within-problem learner. An append-only memory stores representations of previously sampled, target-verified rollouts and their outcome scores; gated cross-attention exposes these records to the draft. Verification distributions provide a distillation objective, while outcome-weighted replay provides an auxiliary learning signal. We distinguish two mechanisms that are often conflated: under exact speculative sampling, adapting the draft changes sampling cost but not the frozen target's conditional output distribution; direct changes to search quality require an explicit draft-dependent selection rule. We give sufficient conditions for exact sampling with adaptive memory, describe replay bookkeeping, and derive a compute ledger and a conditional break-even criterion. We specify a matched-budget evaluation on mathematical reasoning, including a decoder-only control given the same rollout information. This paper presents a method and analysis with an evaluation protocol; empirical improvements in accuracy, acceptance length, and architectural efficiency remain untested.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.