acceptodds
Under review as a conference paper at ICLR 2027

Typed-Reward-Aware Credit Estimation and Routing for Multi-Reward Agentic Reinforcement Learning

Abstract

Practical language-model agents face diverse, task-dependent reward requirements. Common examples concern task outcomes, intermediate processes, and conditions on valid interaction. Reinforcement learning for these agents therefore faces a multi-reward credit-assignment problem, extending beyond distributing a final outcome score across turns. We identify three coupled difficulties: shared temporal propagation can let dense process feedback overwhelm outcome credit; mixed action-type statistics can encourage easily rewarded behaviors; and additive reward fusion can allow task gains to offset violated prerequisites. We propose **TRACER**, Typed-Reward-Aware Credit Estimation and Routing, to preserve these distinctions throughout credit assignment. TRACER estimates reward-specific advantages with separate horizons, baselines, and normalization scales, refines comparisons by action type where appropriate, and uses conditional rewards to gate which pathways enter each update. The normalized pathways are then fused into a single policy advantage. We instantiate three common reward types—outcome, process, and conditional rewards—in search-augmented question answering. Experiments across seven benchmarks and three backbones show improvements in answer accuracy over group-relative baselines. Component analyses connect these gains to auxiliary-reward horizons, action-type comparisons, and sustained format compliance. Together, these results support treating temporal responsibility, action comparability, and conditional eligibility as joint requirements for credit assignment in practical multi-turn agent RL.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.