ZeroHero: Training Small Language Agents from Scratch
Abstract
Large pretrained language models provide capable policies for interactive tasks, but how much model capacity is necessary for a specialized language agent? We introduce ZeroHero, a framework for training small language agents from scratch through expert-trajectory pretraining and adaptive post-training. A central challenge is deciding how to allocate training between expert supervision and the agent's own experience as its competence changes. We introduce Bandit-based Adaptive Training Selection (BATS), which uses a contextual bandit to select blocks of DAgger-style imitation learning or PPO updates based on measured learning progress. We study this approach under Reinforcement Learning with Oracle Dialogue (RLOD), where an agent combines environment actions with natural-language questions to an oracle that provides otherwise hidden task information. We construct dialogue-augmented ALFWorld household tasks and Craftax navigation tasks, and train 5–10M-parameter executors from random initialization, first on successful expert trajectories and then through online interaction. The expert supplies training supervision, while an LLM oracle remains available during execution. On ALFWorld, BATS achieves 89.6% success, compared with 67.9% for the expert, and outperforms fixed PPO, DAgger, and mixed-loss training from the same starting checkpoint. BATS also supports recovery from weak starting policies where pure PPO collapses. Across environments and starting policies, it selects different training strategies, ranging from predominantly reinforcement learning to sustained expert-guided imitation. These results show that compact language agents can surpass their demonstration expert and highlight the importance of adapting training regimes to the task and the learner's competence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.