acceptodds
Under review as a conference paper at ICLR 2027

PolicyTTT: The Policy as the Product of Test-Time Training for Clinical Rubric Generation

Abstract

Test-time training (TTT) improves solution discovery by updating a model during search at test time. However, methods such as TTT-Discover are primarily evaluated by the best solution found, leaving unclear whether the adapted policy itself retains these gains. We study this question in clinical rubric generation and find that discovering a high-quality rubric does not necessarily enable the adapted policy to reproduce its quality when generating anew from the same conversation alone. We introduce PolicyTTT, which explicitly trains the policy to retain improvements found during search. The policy generates rubric-construction procedures that a frozen executor follows to produce rubrics. Its key component is a bridge that trains the policy to generate successfully refined procedures from the conversation alone, without search history or evaluation feedback, connecting search-time improvements to fresh generation after adaptation. Across seven HealthBench themes, with one adaptation conversation per theme, PolicyTTT improves fresh-generation criterion recall over the base policy on all seven conversations and over TTT-Discover on six. Improvements also extend to search itself: PolicyTTT matches or exceeds TTT-Discover's best discovered rubric on six conversations. On the communication conversation, removing the bridge reduces fresh recall from to and the best discovered rubric's from to . These results show that discovering a strong solution does not by itself ensure that search gains are retained in the policy, and that explicitly training for retention can improve both fresh generation and search.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.