acceptodds
Under review as a conference paper at ICLR 2027

Bi-Level Adaptive Regulation of Intra- and Inter-Task Learning for Compact Long-Horizon Agents

Abstract

Developing agents that autonomously complete long-horizon tasks is an important goal in AI research. Repeated calls to large-scale language models during long-horizon interactions can incur high inference costs, motivating the development of capable agents through targeted training of compact language models. However, existing training paradigms provide limited support for adaptive regulation within and across tasks. At the intra-task level, outcome rewards or undifferentiated process rewards offer limited guidance for switching among reasoning modes. At the inter-task level, fixed task distributions can become misaligned with agents' evolving capabilities. As a result, compact agents may persist with unsuitable reasoning modes and repeat ineffective actions, increasing interaction costs without effectively advancing the task. To this end, this paper proposes ARETE (Adaptive REgulation with Tailored Experiences), a long-horizon reinforcement learning framework that combines intra-task reasoning regulation with inter-task practice regulation. For intra-task regulation, inspired by the self-regulated learning (SRL) framework from educational psychology, ARETE organizes reasoning into three modes: forethought, performance, and self-reflection, using differentiated, verifiable feedback to regulate their use. For inter-task regulation, inspired by the zone of proximal development from developmental and educational psychology, ARETE generates verified scaffolded tasks and compares unaided and scaffolded performance to estimate learning potential, guiding capability-matched task generation and practice allocation. Experiments on two complex interactive benchmarks show that ARETE improves compact agents, enabling a 2B-parameter model to perform competitively with models exceeding 100B parameters.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.