Structured Self-feedback In-Context Reinforcement Learning for Test-Time Adaptation
Abstract
Unsupervised test-time self-improvement remains challenging because LLMs must identify and reuse reliable experience from their own generations without access to ground-truth or external verifiers. In-context reinforcement learning (ICRL) provides a training-free mechanism for test-time adaptation by accumulating self-generated responses and feedback in context to improve reasoning performance. However, existing ICRL methods typically derive pseudo-labels through majority voting, which treats generated responses as an unordered set and overlooks the structured evidence provided by different reasoning strategies. As a result, answers supported by only a narrow subset of reasoning processes may be indistinguishable from those consistently supported across diverse strategies. In this paper, we introduce Structured Self-feedback In-Context Reinforcement Learning (SS-ICRL), which constructs structured self-feedback from multi-strategy generations for unsupervised test-time self-improvement. Specifically, we propose a Reliability-Hill (RH) score that jointly models within-strategy agreement and cross-strategy support, enabling effective pseudo-label selection based on structured reasoning evidence. To gudide self-improvement, SS-ICRL further integrates the self-feedback into strategy-specific contextual experience across optimization rounds. Experiments across multiple mathematical reasoning benchmarks show that SS-ICRL consistently improves reasoning performance in unsupervised test-time settings without updating model parameters.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.