Dr.Q: Multi-granularity Credit Decomposition For Long-Horizon Tool-Using Agents
Abstract
Long-horizon tool-using LLM agents fail for reasons that final rewards rarely localize. We observe that such trajectories have a natural recursive structure: a rollout is a sequence of segments, and each segment contains a sequence of states. Credit assignment at each granularity level requires different mechanisms, and—critically—if any granularity level’s local reward is constant within its comparison group, the policy stops decomposing the task at that level. We formalize this as the granularity degeneracy principle and show it is not caught by existing dynamic-sampling filters (e.g., DAPO). We prove that the Anchor’s cross-trajectory comparison is biased when prefixes differ (Proposition 1), and that constant local reward at any granularity causes the policy to stop exploring that granularity (Proposition 2). The latter is a general property of group-relative methods, not specific to DR.Q. On τ-bench, Qwen3.5-9B-DR.Q improves average success from 27.57 to 40.26 over GRPO (5 seeds, ±2.1). Removal ablations show the Rubric contributes +4.8 points, the Anchor +3.2 points, and the Router +3.3 points. We demonstrate granularity degeneracy empirically: in a partially covered environment, the policy stops exploring tools with constant-zero local reward, reducing execution accuracy by 9.4 points. Code is available at https://anonymous.4open.science/r/DRQ-Research-4599.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.