acceptodds
Under review as a conference paper at ICLR 2027

Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised Reinforcement Learning in LLM Reasoners

Abstract

Dense process supervision supplies local feedback but leaves open how that feedback should become a policy advantage. We introduce **PASS** (*Process Advantage Signal Shaping*), an advantage-construction pipeline for group-relative policy optimization (GRPO). PASS combines separate reward-channel normalization and fusion, signal-defined chunks for credit accumulation, and normalization by the remaining chunk count. This design connects the scale of local feedback to the resolution and horizon over which it contributes to a policy update. We evaluate PASS using learned process rewards for mathematical reasoning and teacher-derived token signals for multi-hop question answering. On seven mathematics benchmarks, PASS improves average pass@1 by points over outcome-only GRPO while reducing average response length by approximately . On multi-hop QA with mean-centered normalization, PASS exceeds outcome-only training by and points for two distillation signals, with longer responses. Across both normalization schemes, PASS achieves the highest mean accuracy among the tested process-signal configurations and remains closest to outcome-only under absolute-maximum normalization. These findings show that channel handling, credit resolution, and accumulation should be considered jointly when incorporating dense supervision into policy optimization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.