acceptodds
Under review as a conference paper at ICLR 2027

EffectSet: A Batch Tool Call Is More Than One Outcome

Abstract

A batch tool call can apply some requested effects, reject others, and leave still others uncertain. Treating that call as one success or failure forces a language agent to reconstruct item-level state from overlapping responses, which can abandon unfinished work or repeat completed effects. We introduce EffectSet, a deterministic runtime representation that folds tool events by target item into settled, verify, retry, and escalate sets, then exposes the residual obligations to the agent. We also introduce PartialCallBench, an executable benchmark with 288 fresh tasks, twelve response grammars, 6,912 target items, and 1,496 model-task decisions from Claude Haiku and GPT-5.6 Luna. On the 120-task main split, exact residual-action accuracy rises from 29.2% with raw transcripts and 72.5% with raw-plus-tagged transcripts to 86.7% with item-grouped canonical histories, then 100% when each item receives a derived state. A fully crossed, tags-only control sharpens the mechanism: grouping changes Haiku and Luna by -8.3 and +4.2 points, whereas derived state reaches 100% for both models, gaining 20.8 and 8.3 points over call-ordered tags. An independently generated 48-task Luna split and four unseen synthetic grammars also reach 100% with EffectSet. More decisively, a 32-task replay grounded in four independently documented public batch protocols raises exact accuracy from 71.9% raw and 75.0% tagged to 100% with the state ledger, a paired gain of 25.0 points over tags. The representation is linear-time, processes 120,000 events in 78.2 ms, and reduces prompt characters by 25.2%. EffectSet turns partial batch execution into explicit state that agents can act on exactly and efficiently.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.