acceptodds
Under review as a conference paper at ICLR 2027

Structured Self-feedback In-Context Reinforcement Learning for Test-Time Adaptation

Abstract

Unsupervised test-time self-improvement remains challenging because LLMs must identify and reuse reliable experience from their own generations without access to ground-truth or external verifiers. In-context reinforcement learning (ICRL) provides a training-free mechanism for test-time adaptation by accumulating self-generated responses and feedback in context to improve reasoning performance. However, existing ICRL methods typically derive pseudo-labels through majority voting, which treats generated responses as an unordered set and overlooks the structured evidence provided by different reasoning strategies. As a result, answers supported by only a narrow subset of reasoning processes may be indistinguishable from those consistently supported across diverse strategies. In this paper, we introduce Structured Self-feedback In-Context Reinforcement Learning (SS-ICRL), which constructs structured self-feedback from multi-strategy generations for unsupervised test-time self-improvement. Specifically, we propose a Reliability-Hill (RH) score that jointly models within-strategy agreement and cross-strategy support, enabling effective pseudo-label selection based on structured reasoning evidence. To gudide self-improvement, SS-ICRL further integrates the self-feedback into strategy-specific contextual experience across optimization rounds. Experiments across multiple mathematical reasoning benchmarks show that SS-ICRL consistently improves reasoning performance in unsupervised test-time settings without updating model parameters.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.