acceptodds
Under review as a conference paper at ICLR 2027

Beyond Token-Average Loss: Rebalancing Control and Binding in Agent Policy Distillation

Abstract

Distilling multi-step agent policies into autoregressive language models requires mastering two coupled decisions: high-level operational control and granular argument binding. Under standard action serialization, control opcodes span only one or two tokens, whereas arguments, identifiers, and syntax dominate sequence length. Consequently, standard token-mean cross-entropy subordinates critical control decisions to argument formatting, leading to brittle failures where a single mispredicted opcode derails an execution trajectory. To address this structural imbalance, our action-group-normalized objective () equalizes supervisory weight between control opcodes and argument bindings at each step. In a controlled factorial study on a shared parameter-binding-adapted Qwen3-4B backbone, we evaluate standard cross-entropy () versus balanced across expert demonstrations and matched replay substitution, strictly matching sample exposure and update budgets across three seeds. During evaluation, student policies operate fully autonomously, generating opcodes and arguments end-to-end without external guidance, policy masks, or runtime heuristics. The balanced objective yields double-digit improvements in closed-loop task success: +12.22 percentage points on replay data (95% task-cluster bootstrap interval: [9.17, 15.42], conditional on observed seeds) and +10.56 points on expert data. Fully autonomous balanced students achieve 94.17% (replay) and 95.00% (expert) mean success, matching the 95.00% point estimate of an external teacher-guided reference. In contrast, expanding offline recovery-state coverage yields no benefit over expert demonstrations under standard objectives, demonstrating that balanced supervisory allocation governs policy competence where offline state diversity alone provides no benefit. Finally, stress tests under joint tool-and-control shifts map the operational boundaries of frozen policies. Code and recorded outcomes are provided in the supplementary materials.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.