LIFT-Bench: Learning Investment for Future Transfer in LLM Agents
Abstract
Language agents can improve through interaction and persistent memory, but these gains incur reasoning and learning costs. As agents acquire reusable knowledge for diverse future tasks, they must decide what to learn under limited resources. Existing benchmarks typically emphasize task performance or per-task cost, providing limited support for jointly evaluating long-term learning gains and resource expenditure from a global perspective. We introduce LIFT-Bench (Learning Investment for Future Transfer) to evaluate allocation of a finite global token budget across competing learning opportunities. It contains 690 learning resources and 329 evaluation tasks across eight domains spanning code, tool use, mathematical reasoning, and local physical laws, organized by reusable knowledge units with explicit support relations and domain-specific verification. Inspection, study, practice, and memory construction share one learning budget; transfer is then evaluated with frozen memory and separate per-task inference allowances. A unified interface supports comparisons across three models and eight frameworks using domain-macro performance and return on learning-token investment. We evaluate three policies: Selective Investment for Future Transfer (SIFT), sequential learning, and no learning.Results distinguish the transfer value of learning expenditures and show that larger budgets need not improve future-task performance. LIFT-Bench measures what agents should learn, how they allocate learning budgets, and whether acquired knowledge benefits future tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.