acceptodds
Under review as a conference paper at ICLR 2027

Pincer: Identifiability-Aware RetroClaim Synthesis for Long-Horizon Agent Task

Abstract

action first and writing the user’s request afterwards. Execution supplies a wit- ness, but back-writing can drop constraints: a request may admit incompatible readings even though its recorded trajectory and answer are fixed. We call this failure non-identifiability and separate it from explicitly permitted alternatives, from task difficulty, and from verifier error. We present PINCER, which audits and repairs synthetic tasks from both directions. Its audit combines an indepen- dent reader ensemble, zero-tool and counterfactual probes, and mutation tests on both schema-derived and naturally occurring near misses, and labels each task as contradicted, uncontradicted, certified on a closed domain, or unknown. Its constructive half, RETROCLAIM, fixes declarative claims over the initial state, closes their admissible assignments, regresses them through typed tool effects into an executable witness, and renders the instruction last. On MCP-Atlas (36 servers, 220 tools), 13.4% of reverse-synthesized tasks are underdetermined ver- sus 3.1% under PINCER–RETROCLAIM; the reader ensemble detects contradic- tions with 82.4% precision and 78.6% recall against adjudicated human labels. Training Qwen3.5-9B on the resulting data reaches 43.6% MCP-Atlas pass, 7.4 points above reverse synthesis at matched token budgets and 6.7 points above a distribution-matched reverse baseline, and improves all four transfer benchmarks. On human-authored compositions requiring 13–24 dependent obligations, teacher gating retains only 27.9% of valid tasks, and keeping validated teacher-unsolved tasks adds 2.2 points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.