acceptodds
Under review as a conference paper at ICLR 2027

PRISM: Mitigating Distillation Collapse in Tool-Use LLMs via Position-Level Target Reliability Masking

Abstract

Self-Distillation Policy Optimization (SDPO) offers an appealing paradigm for aligning lightweight open-source language models (<=1B parameters) by matching the token-level distribution of a same-family self-teacher. However, when applied to structured agentic tasks such as tool use, SDPO exhibits a pervasive failure mode: late-training collapse, where policy capability peaks early and degrades as distillation continues. In this work, we uncover a critical pathology behind this collapse: structured tool-calling sequences feature an intrinsic non-homogeneous supervision topology, where closed-set syntax and tool identifiers enjoy near-certain teacher guidance (p approx 1.0, H approx 0), whereas open-vocabulary argument values suffer from diffuse predictive uncertainty (sigma_H = 1.4091, with entropy spikes exceeding 7.5 nats). Diagnostic probes on representation gradients reveal that uncalibrated argument positions exhibit large gradient dispersion (3.51x norm dominance) and lie at wide subspace angles (cos approx +0.1503) relative to structural updates, introducing transverse drift into shared representations during autoregressive training. Furthermore, we show that conventional sample-level filtering is structurally insufficient when reliable and unreliable supervision co-exist within the same sequence (100% co-occurrence in our diagnostic testbed), forcing an intractable compromise between structural starvation and parameter noise leakage. To resolve this dilemma, we introduce PRISM (Position-level Reliability Insulation and Selective Masking), a lightweight framework that operates directly at the position level without requiring any external teacher models: By evaluating continuous predictive confidence on the self-teacher distribution, PRISM constructs a dynamic binary gradient gate that surgically excises volatile, high-variance parameter updates while preserving the large majority of structural supervision: at tau_pos = 0.5 it retains 97.05% of STRUCT syntax delimiters, 96.88% of TOOL_NAME identifiers, and 88.17% of ARG_NAME schema keys. On controlled 120-step multi-seed evaluations across 14 competitive methods and ablation paradigms on real 4xH100 hardware, standalone PRISM achieves the global highest final action accuracy of 88.7% (+3.4 pp over vanilla SDPO, +5.9 pp over sample-level filtering, +15.7 pp over on-policy GRPO) with an outstanding 97.8% peak capability retention, zero external teacher queries, and smooth parameter convergence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.