CARET: Evidence-Grounded Compilation of Verifiable Agent Tasks
Abstract
Engineering workflow traces capture task requirements, tool interactions, and observed workspace changes. Turning these traces into reusable training tasks requires recovering missing context and reconstructing an unresolved state while preserving the source objective. We introduce CARET, an evidence-grounded framework for compiling these traces into executable agent tasks. A shared evidence contract links source requirements, workspace facts, unresolved dependencies, and verifier properties. It guides the joint construction of the initial workspace, instruction, verifier, and reference solution. Admission checks evaluate the initial state, reference repair, and a designated surface-only repair in separate fresh sandboxes. Execution feedback guides bounded revisions, and revised candidates must pass renewed checks before admission. We construct a corpus of 10,139 tasks and collect 24,187 trajectories. The task format supports demonstration collection for supervised fine-tuning (SFT) and resettable interaction for reinforcement learning (RL), with disjoint source-group partitions. Experiments compare CARET with OpenThoughts-Agent and ClawGym at 4B-Instruct, 8B, and 14B scales after SFT and subsequent RL. CARET achieves the highest PinchBench and WildClawBench pure-text scores in all six scale–stage comparisons. After RL, its margins over the strongest baseline at each scale are 2.00–2.40 and 2.40–2.70 percentage points, respectively. Component ablations support the training utility of the evidence contract, joint construction, and execution-guided revision
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.