Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised Reinforcement Learning in LLM Reasoners
Abstract
Dense process supervision supplies local feedback but leaves open how that feedback should become a policy advantage. We introduce **PASS** (*Process Advantage Signal Shaping*), an advantage-construction pipeline for group-relative policy optimization (GRPO). PASS combines separate reward-channel normalization and fusion, signal-defined chunks for credit accumulation, and normalization by the remaining chunk count. This design connects the scale of local feedback to the resolution and horizon over which it contributes to a policy update. We evaluate PASS using learned process rewards for mathematical reasoning and teacher-derived token signals for multi-hop question answering. On seven mathematics benchmarks, PASS improves average pass@1 by points over outcome-only GRPO while reducing average response length by approximately . On multi-hop QA with mean-centered normalization, PASS exceeds outcome-only training by and points for two distillation signals, with longer responses. Across both normalization schemes, PASS achieves the highest mean accuracy among the tested process-signal configurations and remains closest to outcome-only under absolute-maximum normalization. These findings show that channel handling, credit resolution, and accumulation should be considered jointly when incorporating dense supervision into policy optimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.