C^3: Rethinking Process Supervision via Confidence-Guided Credit Correction in Multi-Turn Reinforcement Learning
Abstract
Large language models are increasingly trained as multi-turn agents that interact with external tools. However, outcome-only reinforcement learning can assign incorrect credit to local bad patterns: when a trajectory succeeds, erroneous behaviors in intermediate turns may still receive positive advantage and be reinforced. To address this issue, we propose Confidence-Guided Credit Correction (C³), a plug-in method for local credit correction. C³ identifies bad-pattern spans and selectively corrects positive credit after advantage estimation, with the correction strength adaptively controlled by the behavior policy's confidence in the corresponding behavior, while leaving the task reward and original credit estimation unchanged. Across multiple multi-turn benchmarks, C³ improves performance under both critic-based and critic-free reinforcement learning methods. Injecting the same behavioral signals as turn-level or outcome rewards does not consistently improve performance. Our results show that local credit correction provides an effective form of process supervision for outcome-driven RL.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.