Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents
Abstract
Long-horizon research requires effective meta-reasoning, i.e., deciding what to investigate next as evidence accumulates. Learning this capability is challenging because the relevant decisions are embedded within long execution traces and their consequences are delayed. We introduce Meta-reasoning for Iterative Research Agents (MIRA), a hierarchical research harness whose outer loop curates context from a persistent Git-based research state and selects the objective for the next inner-loop execution agent. MIRA improves how models allocate additional inference: on IMOProofBench Advanced, it raises GPT-5.5's performance from 67.1% to 100% and helps it explore and pivot to new proof directions; in neural-architecture hill-climbing, it discovers improvements to both Residual Matrix and Loop Transformers. Isolating these decision boundaries also enables targeted credit assignment. We pretrain the language model as a generative critic that forecasts expected remaining return from partial research states across environments, then use the same model as both actor and its own critic in single-rollout asynchronous policy optimization over meta-reasoning decisions. We call this shared model MIRA-AC. By jointly training Qwen3.6-27B-Instruct to value the future potential of a research state and decide what to investigate next, while focusing policy-gradient updates on decisions comprising roughly 25% of generated output, MIRA-AC achieves gains of up to 10–15% across several autoresearch environments. It does so while training on readily available public hill-climbing signals, learning from its own research through proxy feedback rather than sparse or costly final-evaluation labels. Overall, MIRA-AC demonstrates that the strategy for conducting research can itself become a learned, reusable policy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.