acceptodds
Under review as a conference paper at ICLR 2027

More Is Not Always Better: Overlapping Gains and Composition Failures in Agent Adaptation

Abstract

Language-model agents can adapt through reflective instructions, retrieved memory, and weight updates. Should combining useful channels improve performance? We study this question with a frozen factorial design, reversed presentation order, and matched token allowances. On 77 held-out coding tasks with Qwen3-8B, 1,386 condition records from 1,232 distinct generations yield no Holm-corrected interaction in either the primary analysis or a grading-replay sensitivity analysis. The primary reflection–memory interaction is percentage points. A task-level audit separates four overlapping-gain cases, where both individual channels and their combination succeed, from three observed combination-damage cases. Adding memory to reflection rescues two tasks and loses two others, leaving the total unchanged. Thus equal scores can hide changes in solved tasks, and a negative interaction estimate alone does not identify destructive interference. An exploratory 270-generation study on reused validation tasks does not establish a benefit from targeted transformation checks. We then test composition decisions on 169 new tasks with 845 generations: a development-fitted, outcome-blind router passes 47 tasks versus 42 for the development-selected fixed combination ( points; Holm ). Reducing memory from three items to one passes 44 versus 42 ( points; Holm ). Evaluation errors remain material: 36 records after replay in the factorial study, 20 in the exploratory study, and 81 in the new-task study. Our evidence distinguishes overlapping benefits from observed composition failures and connects diagnosis to prospective selection tests, while leaving robust antagonism, ordering advantages, and reliable adaptive composition unconfirmed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.