acceptodds
Under review as a conference paper at ICLR 2027

Reusing Trainable States, Not Just Solutions: Policy-Archive Search at Test Time for Automated Discovery

Abstract

Test-time discovery can accumulate two coupled histories: the solutions found during search and the trainable policies produced by successive test-time updates. Test-time search with frozen language models reuses discovered solutions, while test-time training additionally updates the model using verifier feedback on newly generated candidates. Yet such methods typically continue training from the most recently updated policy; earlier policies are never again eligible for an update. Solution reuse preserves sampled outcomes, but not the earlier proposal distributions that produced them. Moreover, a policy's proposal behavior depends on the evolving context provided by previously discovered solutions—the Solution Archive. A later Solution Archive can therefore elicit different behavior from a frozen historical policy. Empirically, re-evaluating policies saved at different training stages under the same later Solution Archive reorders their relative performance—earlier states can even outperform the latest one. The starting point of each update is therefore a search decision, not a fixed sequence. We introduce PAST (Policy-Archive Search at Test Time), which extends reuse from solutions to trainable policies. Its Policy Archive allows learning to resume from a historical state under the current Solution Archive, creating a new continuation branch. We instantiate PAST with ProgressEMA, which uses observed capability and delayed parent-to-child progress to choose which archived policy receives the next update. Across three verifier-driven mathematical discovery tasks, PAST–ProgressEMA reaches better frontiers than the Current-only method, which always continues from the latest policy, while maintaining high per-epoch proposal quality even when the cumulative frontier plateaus. These results show that reusing trainable policies, not just solutions, can improve test-time discovery. Code is available at https://anonymous.4open.science/r/submission-artifact-105E/README.md.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.