acceptodds
Under review as a conference paper at ICLR 2027

TORCH: Jointly Evolving Task Execution and Online Supervision

Abstract

Language-model agents increasingly perform complex tasks through sustained interactions with tools and environments. During execution, agents may repeat ineffective actions, overlook contradictory evidence, or terminate before adequately verifying results. Existing methods primarily use post-hoc reflection to improve subsequent attempts or online supervision to correct ongoing execution, with some further retaining skills and memory for experience reuse. However, task execution and supervisory assessment involve distinct decisions, and how to accumulate their respective experience from the same interaction so that both roles can continually improve remains an open question. In this paper, we propose TORCH, a framework that corrects execution trajectories through online supervision and supports separate experience accumulation for the worker and supervisor. Its core idea is to maintain independent long-term memories for the two roles, enabling the worker to learn task-solving knowledge and the supervisor to learn diagnostic and guidance strategies, so that both roles can continually evolve while model parameters remain fixed. Further, we introduce TORCH-GA, a harness built on GenericAgent that fully implements the TORCH framework. We evaluate runtime supervision on external harnesses and memory evolution and transfer in TORCH-GA. Empirical results demonstrate that TORCH improves task performance, that joint evolution achieves higher mean later-round performance than either one-sided variant across three independent ten-round runs per condition, and that accumulated experience remains useful across models and task domains without further memory updates. Our code is available at https://anonymous.4open.science/r/TORCH-D8FF/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.