acceptodds
Under review as a conference paper at ICLR 2027

Replay-Verified Factor-Level Distillation for Tool-Trace Repair

Abstract

Distilling tool use from a black-box teacher is complicated by a simple failure mode: a repair can be fluent and schema-valid yet remain incorrect under execution. We introduce replay-verified factor-level distillation (RVFD), a procedure that treats teacher outputs as proposals rather than ground truth. RVFD decomposes a proposed repair into schema-aligned factors, applies each edit to the failed trace, and retains an edit as supervision only when replay reaches the desired outcome. On BFCL-CF, full-parameter Qwen2.5-7B and 14B students trained with RVFD improve greedy repair success over direct teacher-output SFT by 0.0547 and 0.0534, respectively, without a teacher or verifier at test time. An augmented factor student reaches 0.9199 greedy success, compared with 0.8893 for whole-trace execution-filtered SFT; this comparison also differs in training-set size. On naturally occurring errors in archived BFCL outputs, local repair raises AST accuracy from 0.5939 to 0.6774, while whole-trace rewriting with beam4 and checker access reaches . Optional test-time replay improves accuracy further, but does not surpass CAR-point or oracle-step CausalFlow on a shared candidate pool. These results support replay as a source of local credit assignment for tool-trace repair, rather than as a general solution to end-to-end agent learning. Code and data are available at https://anonymous.4open.science/r/rvfd-aaai27/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.