Class Semantics Are Residual Supervision: Breadth First, Depth on Demand
Abstract
Many language-model applications repeatedly make decisions over a fixed output space rather than generate open-ended text, such as intent classification, request routing, and relation classification. These systems often carry an explicit semantic specification of that decision space alongside each input, including class names, definitions, and rules that distinguish neighboring labels. This raises a basic question: how much of that specification must remain explicit once the model has learned part of the decision space? We show that class semantics behave as residual supervision: their value is determined by the boundaries that the model and labeled examples have not yet resolved. As supervision and capability resolve more of these boundaries, the value of explicit semantics decays, and what remains concentrates on the still-unresolved boundaries. On CLINC150, the gain of full descriptions over class names falls from about +2.5 Macro-F1 points with five examples per class to zero with one hundred. Catalog design then becomes a semantic-allocation problem: how much depth each class should carry, and where the remaining depth should be paid for. Per-class benefit rankings change across training seeds, which makes depth-first allocation unreliable at small budgets. This instability motivates a more reliable allocation order in the regimes we study: breadth first, depth on demand. Keep a 2–4 word gloss for every class, then restore full descriptions only at unresolved boundaries. The two-pass verified compact catalog uses 40% of the full-catalog tokens and recovers 84% of its gain; the hybrid uses 58% of the tokens and bounds the loss to the full catalog at 0.4 Macro-F1 points (95% one-sided, ten seeds). Token-matched truncation and retrieval fall back to names-only performance. The remaining burden need not stay in the prompt: continuing training under the deployment prompt moves most of the catalog dependence into the model weights, recovering 97% of the accuracy lost when the catalog is removed, while a tenth of that budget recovers 11%. Catalog dependence also carries a seven-point format component that does not track the description gain. Explicit semantics are a budgeted resource: measure what remains, cover the decision space before adding depth, and move stable semantics into the weights when retraining is available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.