Mem-C: Scaling and Allocating Credit Beyond Dense Rewards in Long-Horizon Agents
Abstract
Long-horizon language agents face severe temporal credit-assignment challenges, as the consequences of intermediate decisions may only emerge many steps later. Dense process supervision and tree-based training alleviate this issue by providing additional downstream feedback, but existing work largely treats credit assignment as a one-dimensional notion of reward density. This overlooks a key fact: the same credit budget can be distributed differently across decisions, and additional supervision may have very different marginal value depending on where it is assigned. We therefore ask: how does agent performance scale with credit, and how should a fixed credit budget be allocated across long-horizon decisions? To answer this question, we introduce **Credit Assignment Geometry**, which characterizes step-level credit along three dimensions: ① ***Q**uantity*, ② ***H**eterogeneity*, and ③ ***U**tility*. We find that increasing credit quantity can yield diminishing returns, while allocations with comparable quantity and coverage can yield substantially different performance due to heterogeneous credit utility. Motivated by these findings, we propose **Mem-C**, which uses branch-outcome dispersion to identify high-utility decisions for importance-aware branch allocation. Under matched credit budgets, **Mem-C** outperforms the evaluated baselines on all three transfer benchmarks, including **12%↑** on HotpotQA. Code is available at [Mem-C](https://anonymous.4open.science/r/Mem-C-543B/).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.