Let the Weak Judge: Annotation-Free Process Rewards for Multimodal E-commerce Relevance
Abstract
Outcome-only reinforcement learning often fails to produce evidence-grounded reasoning in discriminative multimodal tasks: with a small label space, GRPO routinely exploits lexical shortcuts and fabricates forced associations to match the correct label. Process reward models (PRMs) address this in principle, but typically require expensive step-level annotations and can mis-score reasoning outside their training distribution. We propose WISE-GRPO, which turns step-level verification into a verifiable answer-recovery test: we truncate the policy’s reasoning at intermediate steps, mask its upfront answer, and ask a committee of weak models to resume reasoning; a step is rewarded only when its explicit evidence guides the committee back to the correct label. The policy adopts an answer-first Post-CoT paradigm to anchor decisions, while the committee uses a Pre-CoT paradigm for independent continuation-based verification. Increasing step weights emphasize the integrative decisions near the end of the task template. Across four e-commerce benchmarks and two multimodal backbones, WISE-GRPO achieves the best accuracy among the evaluated methods in all eight settings. Its gains over outcome-only GRPO range from to accuracy points, and further analysis shows it reduces forced-association behavior, without requiring step-level human annotation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.