PolicyTTT: The Policy as the Product of Test-Time Training for Clinical Rubric Generation
Abstract
Test-time training (TTT) improves solution discovery by updating a model during search at test time. However, methods such as TTT-Discover are primarily evaluated by the best solution found, leaving unclear whether the adapted policy itself retains these gains. We study this question in clinical rubric generation and find that discovering a high-quality rubric does not necessarily enable the adapted policy to reproduce its quality when generating anew from the same conversation alone. We introduce PolicyTTT, which explicitly trains the policy to retain improvements found during search. The policy generates rubric-construction procedures that a frozen executor follows to produce rubrics. Its key component is a bridge that trains the policy to generate successfully refined procedures from the conversation alone, without search history or evaluation feedback, connecting search-time improvements to fresh generation after adaptation. Across seven HealthBench themes, with one adaptation conversation per theme, PolicyTTT improves fresh-generation criterion recall over the base policy on all seven conversations and over TTT-Discover on six. Improvements also extend to search itself: PolicyTTT matches or exceeds TTT-Discover's best discovered rubric on six conversations. On the communication conversation, removing the bridge reduces fresh recall from to and the best discovered rubric's from to . These results show that discovering a strong solution does not by itself ensure that search gains are retained in the policy, and that explicitly training for retention can improve both fresh generation and search.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.