acceptodds
Under review as a conference paper at ICLR 2027

Test-time Reinforcement Learning in Imperfect Information Games

Abstract

Test-time reasoning has significantly improved performance in domains ranging from games to language models. However, principled test-time policy improvement remains a challenge in two-player zero-sum imperfect-information games. Existing solutions are limited to tabular methods or single gradient step updates. In this work, we investigate policy-gradient algorithms as a method for scalable test-time reasoning. We extend the concept of gadget games, a tabular technique for test-time search, to the reinforcement learning setting. Unlike prior approaches, we represent the gadget game implicitly by training a state sampler and virtual actor policy rather than explicitly by constructing the gadget, thereby removing constraints on subgame size. Furthermore, we prove that a fixed regularization policy bounds the strategy change within a subgame, making simpler Bayesian subgame solving viable without safe gadgets. In small games, test-time training substantially improves weak blueprints, while strong blueprints may degrade slightly. In large games, it improves over the blueprint in head-to-head play and generally outperforms the update-equivalence framework.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.