Learning to Reason from Feedback: Consistency-Aware Post-Training for Language Models
Abstract
Post-training has made substantial progress at optimizing the final outcome of reasoning tasks, but an equally important ability is rarely optimized as a target in its own right: how a model integrates newly acquired feedback into its subsequent reasoning and decisions. This becomes acute once the model must interact with external tools and execution environments, where it has to judge when additional feedback is warranted, absorb new information without derailing its existing line of reasoning, and tell which interactions actually contributed to success. We treat reasoning from feedback as an explicit post-training objective. To this end we introduce three complementary training signals: a contrastive representational constraint that keeps the model's reasoning state semantically coherent across an interaction while it absorbs the new observation; an uncertainty-based interaction policy that leads the model to seek feedback selectively rather than reflexively; and process-level credit assignment from counterfactual trajectory value estimation, which separates genuinely helpful interactions from redundant or harmful ones. The three are optimized jointly with Group Relative Policy Optimization (GRPO) and trained end-to-end on trajectories the model generates through interaction with a real execution environment. Instantiated in a sandboxed code-execution environment and evaluated on AIME 2024/2025, MATH-500, LiveCodeBench, and GPQA Diamond, our 7B model improves over the strongest same-size agentic RL baseline by 4.8 points on average. Analysis shows that the representational constraint improves feedback-utilization precision by 23%, while uncertainty-based gating reduces redundant interactions by 41%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.