CostDescent: Decoupling Cost From Rewards for Efficient Tool-Integrated Reasoning
Abstract
Tool-integrated reasoning (TIR) expands the capabilities of language models but incurs substantial token and tool-use costs. Efficient TIR therefore seeks to retain task performance at reduced cost. Existing approaches often incorporate resource cost into the reward, so that lower cost can be rewarded at the expense of task success. Subsequent refinements temper this interaction but still keep both in one learning signal. Drawing on goal-conditioned RL (GCRL), we instead let task reward alone evaluate outcome, while resource cost enters policy optimization as a goal input. We introduce , a goal-conditioned training framework that expresses these goals relative to a task-dependent policy-induced reference. From on-policy performance-cost outcomes, constructs a task-aware target reference favoring lower-cost behavior, relabels achieved costs around it, and applies a goal-routed update. We establish sufficient conditions under which this update induces local cost descent. Across six mathematical-reasoning and knowledge-intensive benchmarks, achieves the strongest overall task performance and tool productivity. These results establish task-aware cost descent as an effective approach to Efficient TIR and suggest pursuing behavioral targets outside the task reward in broader LLM reinforcement learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.