acceptodds
Under review as a conference paper at ICLR 2027

Auditing Inferred Tool-Call Dependencies as Proxies for Execution Impact

Abstract

When dependency violations are used to score reordered agent trajectories, inferred edges serve as proxies for execution impact. We audit this assumption through fixed-call replay, swapping two recorded tool calls while preserving their identities and arguments and measuring changes in observations, final state, and reward. Across more than transpositions on -bench and BFCL, we compare token overlap with typed exact value matching between outputs and arguments and with a published typed-ID extractor. On the same eligible pairs, exploratory analyses show that typed matching improves prediction of observation changes relative to token overlap, with Matthews correlation coefficients of –. The published extractor reaches –. A predictor based on whether the originally later call changes state during baseline or prefix execution reaches –, exceeding typed matching and the published extractor in every evaluated setting. This predictor requires environment replay and additional per-call state access, but no pairwise transposition sweep. The advantage extends to trajectory-level observation-change scoring, although gains over simple inversion counts remain uncertain in two of six arms when response positions are fixed. At a fixed inspection budget, it increases observation-change yield but lowers state-change yield relative to typed matching. Fixed-call replay measures execution with supplied arguments; it does not establish whether the agent could generate the reordered calls or identify direct causal dependencies. These findings motivate outcome-specific validation of dependency-based trajectory scores.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.