When Interaction Is the Bottleneck: Evaluating and Scaling Agents Under Costly Feedback
Abstract
Large language model (LLM) agents rely on external feedback to test hypotheses, retrieve information, and verify answers. When feedback is costly, capabilities depend on how they acquire information as well as how they reason over it. We introduce ExpGYM, an evaluation suite built on pre-cached feedback and deterministic surrogates for hyperparameter tuning, multi-hop search, and evidence audit. ExPGYM enables repeated evaluation under different feedback budgets without rerunning costly external work, while preserving feedback charges. Moving from Free Feedback to Tight Budget reduces overall performance across all 6 models and changes task-level model preferences. Gemini 3.8 Flash remains the strongest model overall, yet its multi-hop search AVG falls from 71.9% to 17.6% between Free Feedback and Tight Budget. On ParamNet's Adult task, selecting the Free Feedback leader yields a Tight Budget Gap Score 15.22 below the best available model. Trajectory analyses show that limited feedback restricts exploration and evidence completion, while additional requests do not always provide new information. We introduce POOLACT to coordinate acquisition through shared observations and exploration state. With 4 agents under matched per-agent budgets, POOLACT raises Qwen3.8's Moderate Budget evidence audit exact evidence-set accuracy from 68.8% with independent rollouts to 93.7%. These findings highlight feedback acquisition in evaluating agents and coordinating groups under limited budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.