Separating Structural Semantics from Inference Allocation in LLM Agents
Abstract
LLM-agent comparisons often change both what the model is asked to do and how much inference it receives, then treat realized resource consumption as a control. We separate structural semantics , ex-ante allocation and realized downstream cost , and explain why adjusting for does not in general estimate the effect at fixed allocation. Our intervention requests separately seeded candidates under byte-identical prompts; in sequential control a fixed plurality rule selects the action. We audit prompt invariance using full-message hashes in program synthesis and system-prompt hashes plus prompt-construction tests in control. In a preregistered factorial over four models, four cooperative-cooking layouts and five seeds, allocation contrasts range from to soups per episode. All four percentile-bootstrap intervals lie below zero, but each admits a drop smaller than our preregistered practical margin, so these data do not establish practical harm. Structural effects remain imprecisely estimated. In a separate matrix of 2,624 HumanEval candidate pools, both the frozen and corrected original-test scorers place the structural effect at within the fixed equivalence margin. This oracle-success metric has no aggregator and is not a replication of the control result. Observed changes run opposite to both preregistered mechanism predictions; the mechanism remains unresolved. The results concern the tested sampling-and-plurality policy, whose candidate count and selection rule are not separately identified.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.