acceptodds
Under review as a conference paper at ICLR 2027

SteerDistill: Quantizing LLM Agents via Online Teacher-Steered Trajectory Distillation

Abstract

Quantization makes multi-turn LLM agents brittle: a single erroneous decision alters every subsequent interaction and can derail the task. We observe that quantization rarely breaks an agent's overall plan; instead, it corrupts a few executable decisions, such as a tool call with the wrong interface or a misgrounded argument. Existing recovery methods either learn from fixed full-precision data that misses the states a quantized agent reaches, or supervise the student's own rollouts without locating the error and rely on later rollouts to avoid it; neither corrects the turn at which a rollout goes wrong. We introduce SteerDistill, a quantization-aware distillation method built on *online teacher steering*. When a quantized student's rollout diverges from a successful full-precision trajectory, the teacher issues a steering prompt at the divergent turn; the student regenerates that turn itself, the environment verifies the correction, and the rollout continues from the corrected state. The steered behavior is then distilled into the student without the prompt, so no teacher is needed at inference. Across five agentic benchmarks and three models with 4-bit weights, \nickname outperforms the strongest quantized baseline by 1.67–2.18 average points and stays within one point of BF16. On -Bench, it also exceeds quantization-aware reinforcement learning with 67% fewer training GPU hours.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.