acceptodds
Under review as a conference paper at ICLR 2027

ONE SEARCH, TWO VALUES: TEST-TIME SEARCH EFFICIENCY IN THREE-PLAYER GENERAL-SUM GAMES

Abstract

With three or more players, a search should back up the value aligned with the players’ objective. Objectives such as win probability or rank credit make that aligned value coarse: a player’s actions are mostly exactly tied or far apart, so a few-simulation search over the aligned value is often left undecided. The cardinal score is dense and decisive but misaligned, since the score-optimal action can be strictly worse in win probability. We carry both values through one search. A tree search guided by the objective-aligned value head also backs up the cardinal head’s vector through the same nodes, so both action-value vectors reach the root with no additional search simulation. A gated readout then plays the aligned decision where the aligned margin clears a calibrated gate threshold and the cardinal ordering where it does not. The readout leaves the tree unchanged, so its effect is confined to the states where it fires, and its expected gain is the firing probability times the accuracy difference of the two readouts on those states, which we measure. On an exactly solved eight-token ranking game and on a reduced three-player Chinese Checkers, the gated readout at 64 simulations per move beats the aligned readout at 64 and 128 simulations and the cardinal search at both budgets. The aligned readout needs about seven times the simulation budget to match it on the ranking game under the rank objective, and on Chinese Checkers it does not match it within 1024 simulations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.