Active Agentic Learning
Abstract
Training capable large language model agents requires costly trajectory annotations that contain multi-step reasoning, tool calls, and environment interactions. Active learning can reduce this burden by prioritizing informative trajectories, but conventional uncertainty scores are poorly matched to agentic data: consequential decisions are diluted by predictable transcript tokens, while environment observations can vary independently of the model's pre-call confidence. We formulate Active Agentic Learning and propose Pseudo-Influence, a trajectory-valuation criterion that directly estimates training impact through pseudo-label perturbation rather than predictive uncertainty. Our efficient estimator provides theoretical guarantees without requiring a convex objective or inverse-Hessian computation. We evaluate Qwen3-8B and Qwen3-32B over five acquisition rounds in two complementary settings: command-line interaction on Terminal-Bench 2.0 and tool use on ACEBench-Agent. Across all 16 evaluation points before full-pool training, the best Pseudo-Influence variant ranks first and exceeds the strongest non-Pseudo-Influence baseline by 1.4–5.3 pass@1 points, with a mean gain of 3.1 points. In the first acquisition round, combining Pseudo-Influence with diversity improves over random selection by 4.8 and 5.5 points on Terminal-Bench 2.0, and by 6.7 points for both model scales on ACEBench-Agent. These results indicate that training-impact valuation and diversity are complementary signals for annotation-efficient agent learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.