Go for the Reward: Monte Carlo Tree Search for Training Sub-Billion Language Model Agents
Abstract
Training sub-billion language model agents without external teachers is difficult when their initial policies rarely complete tasks and receive little guidance from task-level failure feedback. We introduce Go for the Reward (GFR), a framework that learns directly from environment interaction without teacher demonstrations or task-specific supervised initialization. Inspired by the search–learning loop of Go-playing AI, GFR alternates training-time Monte Carlo tree search with policy learning. Search explores alternative continuations from shared interaction prefixes to discover successful trajectories. Execution constraints complement these trajectories with feedback on malformed or inadmissible actions and redundant interactions, including in unsuccessful attempts. The policy learns from this experience through successful-trajectory supervision, outcome-supported sibling-branch comparisons, and execution penalties, and then guides subsequent search. Starting from pretrained Qwen3-0.6B, GFR achieves 39.55% unseen success on ALFWorld and a WebShop Task Score of 59.61 without evaluation-time search. Both results are the highest among the compared teacher-free methods. On ALFWorld, this teacher-free training reaches performance comparable to teacher-assisted supervised fine-tuning followed by reinforcement learning. GFR also learns more efficient execution paths, reducing environment interactions by about 31% relative to its no-tree variant on a shared set of successfully completed WebShop tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.