acceptodds
Under review as a conference paper at ICLR 2027

SIPA: Test-Time Reinforcement Learning through Speculative Policy Adaptation

Abstract

Reinforcement learning (RL) has become a central paradigm for language-model post-training, yet its gains do not always transfer to unfamiliar deployment environments. Experience-based test-time adaptation has emerged as a promising approach, drawing on accumulated experience to guide subsequent decisions. However, collecting relevant experience is computationally demanding, and the resulting guidance may be inaccurate or poorly matched to the current state. We introduce Speculative Inference-Time Policy Adaptation (SIPA), a test-time RL framework that adapts the current action distribution using immediate execution feedback, without parameter updates or pre-collected trajectories. While the actor reasons, a lightweight branch proposes and tests candidate actions in isolated copies of the current environment state. Asymmetric Policy Adaptation (APA) combines the actor's selection preferences with the critic's execution-based rejection feedback. A KL-regularized logit update incorporates these signals before action commitment while bounding deviation from the actor's prior. We establish stability bounds under feedback error and sufficient conditions for expected-return improvement. Across five multi-turn benchmarks, SIPA achieves a 22.1% relative improvement in average accuracy over the strongest baseline using only 69.6% as many committed steps.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.