acceptodds
Under review as a conference paper at ICLR 2027

Hard Negatives Should Match the Student: Teacher–Student–Judge Preference Distillation for Agentic Tool Use

Abstract

Open-weight language models lag behind frontier systems on "agentic" workflows: long-horizon planning, strict tool-schema adherence, and policy compliance under adversarial user pressure. We present the Teacher–Student–Judge (TSJ) pipeline, which synthesizes Direct Preference Optimization (DPO) pairs by letting the target model (Student) compete against a stronger Teacher inside simulated multi-turn conversations, with a Judge that selects winners and synthesizes hard negatives against the Student's own successful replies. Fine-tuning Qwen3-30B-A3B on 7,747 TSJ pairs improves Tau-Bench Airline from 32% to 44% and Retail from 54% to 61%; general capabilities remain within 1.3 points across four evaluations (MMLU, GSM8K, and HumanEval improve, while IFEval declines 0.9 points). Attribution: four-condition ablations on Qwen3-4B associate removing the Judge-synthesized-negative slice with a 2.2-point lower macro-average; size-matched control scores 1.0 point above distillation-only. Generality and its boundary: the recipe fails when Granite-4.1-8B is trained on Qwen-generated data, but succeeds (+5.8 points, +21% relative) when regenerated with Granite as Student. At identical dataset size, Granite-matched synthesized negatives beat negative-free distillation by +4.1 points. Offline vs. online: at matched scale, offline TSJ-DPO matches online GRPO trained directly in the executable Tau environment (0.432 vs. 0.419 macro-average). These results position student-anchored preference pairs, rather than generic teacher demonstrations, as the highest-leverage data for agentic preference optimization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.