Can Weak Models Explore? Scope Then Cover for LLM Agent Adaptation
Abstract
Adapting a language-model agent to a new environment requires exploration, and the prevailing assumption is that stronger models explore better. We examine two regimes where model capability alone is insufficient for discovery: when the available information does not help locate the target, and when the discovery advantage over uniform inspection diminishes as the candidate space grows. These observations motivate combining task-informed guidance with broad inspection: a strong actor selects a candidate scope, and weak scouts inspect candidates within it in parallel. We instantiate this scope-then-cover decomposition as a single explore action for frozen test-time actors. On ALFWorld, WebShop, and SWE-Bench-Verified, adding the explore action improves task performance over the actor-only baseline by assigning additional inspections to weak scouts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.