acceptodds
Under review as a conference paper at ICLR 2027

The Filter Outweighs the Channel: Decomposing Why Tool-Using Agents Corrupt Structured Records at Scale

Abstract

When a tool-using agent answers a follow-up question over structured records, two jobs sit between the tool's output and the user: selecting the rows the user asked for, and delivering them. Deployed agents differ on both at once, so prior comparisons measured a joint effect. We unbundle them with four conditions over identical frozen payloads from a deployed travel agent: filtering done in-head by the model versus by a deterministic evaluator, crossed with delivery by payload (the model re-types the rows), by reference (the model names the rows and code expands them), or by a typed side channel (rows never pass through the model). The axes are far from symmetric. In flight search, taking the filter off the model recovers most of the collapse (10% to 78% turn fidelity at 100-row tables) and is the only axis with a detectable size dependence; delivery by payload costs about 23 points there, with no size dependence detectable at this resolution. Failure analysis shows why: in-head filtering is a fused interpretation-and-selection act whose set membership degrades with table size, while re-emission drops rows. A replication on the same deployed system's hotel search, a second record domain, transfers the filter-axis collapse (the filter-axis contrast is -23 points per unit of ln K on flights and -21 on hotels) and the typed channel's advantage, but the payload cost does not stay constant there: at 100 hotel rows re-emission falls to 46%, through missing emissions, empty tables and wrong row sets. The deployable rule follows: taking the filter off the model is necessary for size-immunity but not sufficient wherever re-emission itself breaks with table size or with the model family; delivery by reference helps a little; the typed side channel is the one arrangement that holds for every family and in both domains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.