Code-Probe-Adapt: Interactive Heuristic Learning for Self-Evolving Agents
Abstract
Existing LLM-based autonomous agents generally follow two paradigms: ReAct- style methods, which interleave dense reasoning with continuous API calls; and code-as-policy methods, which synthesize executable code to bypass LLM infer- ence at runtime. However, in long-horizon interactive tasks, ReAct agents incur prohibitive inference costs and context overload, whereas offline policies lack test- time adaptability to unforeseen environmental dynamics. To bridge this gap, we introduce Code-Probe-Adapt (COPA), an interactive heuristic learning framework that transitions LLM agents from post-hoc log analyzers to active, online policy designers. Equipped with a persistent REPL (Read-Eval-Print Loop) harness, COPA enables live, on-demand introspection. To enable adaptive code-as-policy, COPA modulates policy execution using adaptive step budgets and event-triggered interruptions, returning control to the agent for immediate diagnostic probing and policy refinement. Evaluated on five games, COPA significantly outperforms re- cent LLM-based policy synthesizers, and demonstrates superior sample efficiency compared to extensively trained reinforcement learning (RL) algorithms. Notably, on the hard-exploration task Montezuma’s Revenge, where offline LLM heuristic learning fails entirely (scoring 0 even after 6.0M steps), COPA’s interactive prob- ing achieves an average score of 4,900 in 2.4M steps. This result also compares favorably against the RL baseline DreamerV3 (1,853.5 at 200M steps, digitized from its learning curve), scoring 2.6× as high with fewer environment interactions. Attained entirely without gradient updates or task-specific engineering, COPA successfully unites the execution efficiency of code with the adaptive power of live debugging.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.