acceptodds
Under review as a conference paper at ICLR 2027

CostDescent: Decoupling Cost From Rewards for Efficient Tool-Integrated Reasoning

Abstract

Tool-integrated reasoning (TIR) expands the capabilities of language models but incurs substantial token and tool-use costs. Efficient TIR therefore seeks to retain task performance at reduced cost. Existing approaches often incorporate resource cost into the reward, so that lower cost can be rewarded at the expense of task success. Subsequent refinements temper this interaction but still keep both in one learning signal. Drawing on goal-conditioned RL (GCRL), we instead let task reward alone evaluate outcome, while resource cost enters policy optimization as a goal input. We introduce , a goal-conditioned training framework that expresses these goals relative to a task-dependent policy-induced reference. From on-policy performance-cost outcomes, constructs a task-aware target reference favoring lower-cost behavior, relabels achieved costs around it, and applies a goal-routed update. We establish sufficient conditions under which this update induces local cost descent. Across six mathematical-reasoning and knowledge-intensive benchmarks, achieves the strongest overall task performance and tool productivity. These results establish task-aware cost descent as an effective approach to Efficient TIR and suggest pursuing behavioral targets outside the task reward in broader LLM reinforcement learning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.