acceptodds
Under review as a conference paper at ICLR 2027

Reinforcement Learning via Structured Deliberation and Fine-Grained Attribution

Abstract

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models by optimizing policies against a single reward signal derived from the final outcome. However, real-world reasoning tasks often involve multiple valid answers or more open-ended objectives that a single outcome signal cannot adequately capture or distinguish. In such settings, outcome-only RLVR assigns a homogeneous advantage to the entire reasoning trajectory, ignoring both the complex reasoning structure and the heterogeneous contributions of individual components. We introduce **s**tructured **o**ptimization via **d**eliberation and **a**ttribution (SODA), a framework that addresses these limitations through structured deliberation and fine-grained attribution. SODA first makes the structure of complex reasoning explicit by decomposing intermediate reasoning into semantically meaningful, locally evaluable decision units, and then attributes credit to each individual unit based on its contribution in the context of the remaining response. These credits are computed via local edits, yielding differentiated training signals without additional policy rollouts or a learned critic. We instantiate SODA on set-valued QA, where each question may have a different number of semantically distinct answers. Reasoning is structured into complementary stages of candidate proposal and selective refinement to produce correct and non-redundant answers. Experiments across multiple datasets and base models show consistent improvements over state-of-the-arts, with ablations demonstrating complementary gains from structured deliberation and fine-grained attribution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.