acceptodds
Under review as a conference paper at ICLR 2027

The Serialization Bottleneck: Separating State Preservation from Property Derivation in Tool-Using Agents

Abstract

Tool-using agents often return intermediate structured objects, such as polygons, graphs, and tables, to the language model as serialized text. The next step then asks the model for two different operations: derive a property that selects the branch, and re-emit the object as an argument of its next tool call. End-to-end benchmarks cannot tell which of the two fails. We isolate them with a controlled diagnostic benchmark over geometry, graph, and tabular pipelines and seven language models. Under standard non-thinking inference, models read surface facts far better (51.3%–99.5%) than they derive global properties (20.8%–49.0%). In multi-step pipelines, calls that must re-emit the full object abort in up to 99.7% of tabular runs. Trace audits show early semantic corruption of the payload, not truncation, and schema-constrained decoding does not prevent it. In the most abort-prone setting (Qwen3-32B, tabular), keeping the full observation but passing an immutable reference as the argument cuts aborts from 99.7% to 1.2%, so the aborts come from re-emitting the payload, not from seeing it. For recurring branch predicates, runtime-computed summaries turn derivation into lookup; predicates outside the schema still need on-demand inspection or code execution. Deriving a property from a structured state and re-emitting that state in the next action are separate failure points, and each needs its own interface fix.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.