Hierarchical Experimentalist Agents
Abstract
Large language models (LLMs) are increasingly used to act in the real world and augment human decision-making, yet most systems rely on parametric knowledge acquired through imitation, optionally improved by post-training, retrieval, or search over fixed data. This paradigm struggles in novel domains or on complex queries that require information unavailable from prior knowledge alone; knowing the laws of physics, for example, does not by itself enable reliable reasoning or long-horizon actions in a complex physical system. We argue that agents therefore need *active experimentation*: the ability to gather targeted query-specific evidence, discover general principles of unseen environments, and convert interaction experience into reusable skills. We introduce **Hierarchical Experimentalist Agents (HExA)**, an in-context, experiment-centric self-improvement framework that (1) iteratively designs and refines query-relevant experiments, (2) incrementally distills experience into reusable and composable skills that accelerate experimentation within and across tasks, and (3) integrates experimental evidence to act or answer queries. HExA can work with any LLM including frontier black-box models, and requires no external supervision, oracle solutions, or offline data. We evaluate HExA on multi-step reasoning, interactive partially observable games, and physics law discovery tasks. To further evaluate when active experimentation is especially necessary, we introduce Interphyre, a set of hard procedural physics tasks where we enable tool-calling and intervention APIs. Current frontier LLMs including GPT-6 Astra fail in these settings: on the most challenging Interphyre tasks, we show that HExA can enable such models to near 100% success. Moreover, HExA improves token efficiency compared to naive agentic and memory baselines and enables generalization via hierarchical skill transfer from easier to harder tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.