acceptodds
Under review as a conference paper at ICLR 2027

Cross-Trajectory Semantic Credit Sharing for Mathematical Reasoning

Abstract

Outcome-supervised methods such as GRPO assign one advantage throughout each response, even when solutions with different final outcomes contain the same local mathematical derivation. Prefix-based branching pools terminal feedback through shared generation histories, yet related local derivations can also arise in independently generated solutions. We introduce Semantic Credit Sharing (SCS), which forms local sharing groups after rollout collection. To match related content expressed in different forms, a frozen encoder represents selected segments using their original text, preceding context, and auxiliary normalized mathematical expressions. SCS clusters these segments within each question, averages the advantages of the distinct responses represented in each cluster, and applies the resulting coefficients to the matched token spans. Other token advantages remain unchanged. Sharing reuses the terminal rewards of the sampled responses. Across three independent training runs on Qwen3-1.7B-Base, SCS improves mean accuracy across five benchmarks by 1.88 percentage points over Dr.GRPO. It also outperforms a within-response assignment-shuffling control by 1.10 points, demonstrating the value of preserving credit-to-segment correspondence. On Qwen3-4B-Base, a single paired run yields a 1.49-percentage-point improvement over Dr.GRPO. These results show that local mathematical content can support feedback sharing beyond common generation histories.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.