acceptodds
Under review as a conference paper at ICLR 2027

Multi-Agent-Inspired Reward Channels for Mathematical Reasoning: A Diagnostic Case Study

Abstract

Multi-agent systems for mathematical reasoning can use separate solver and verifier models. We examine what is retained when answer-level verification and output monitoring are instead encoded as deterministic reward components for a single policy. The five components check correctness, answer agreement, numeric range, behaviour, and final-answer format. On 200 olympiad-level problems, DeepSeek-Math-7B-Instruct achieves 8.5% pass@1, compared with 7.0% after GRPO alone, 4.5% after supervised fine-tuning (SFT), and 2.0% after SFT followed by GRPO. In sparse SFT+GRPO logs, the binary format reward reaches its penalty floor in recorded batches while mean completion length approaches the 512-token limit. A reward component that is constant across a sampled group cancels from the group-relative advantage even when aggregate reward varies. The logs do not establish how often this occurred throughout training or whether the generation limit caused the accuracy decline. This case motivates measuring within-group variation for each reward component.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.