CREDO: Evidence-Conditioned Credit Assignment for Temporal Reasoning Agents
Abstract
Temporal Knowledge Graph Question Answering (TKGQA) requires language model agents to retrieve evidence, select relevant facts, and perform multi-step temporal reasoning. However, the trajectory-level outcome signal in GRPO does not explicitly distinguish the quality of intermediate decisions: downstream failures can penalize useful evidence acquisition, while identical rewards in all-failed groups eliminate the relative outcome signal despite differences in intermediate progress. We propose Credit Redistribution with Evidence-Driven Opportunities (CREDO), which uses intermediate evidence to recover local supervision and determine where outcome supervision acts. CREDO assigns positive stage-level credit through comparisons among rollouts of the same question and uses retained answer overlap as an opportunity proxy to shift part of the outcome-policy weight toward reasoning and answering. The framework preserves terminal rewards and trajectory-level advantages, while redistribution conserves total outcome-policy weight per trajectory. Following the public evaluation protocol of prior work, CREDO achieves 82.4% overall accuracy on MultiTQ, surpassing the strongest published baseline in our comparison by 4.4 percentage points overall and 10.6 percentage points on Multiple questions. On TimelineKGQA, it improves cross-KG accuracy over our reproduced baseline by 5.5 percentage points via direct transfer to ICEWS Actor under matched training and evaluation settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.