acceptodds
Under review as a conference paper at ICLR 2027

Learning to Defer When Routing Changes Competence

Abstract

Learning-to-defer systems assign inputs according to the current competence of the available models or experts. When models train only on the inputs they handle, routing also determines the training exposure that sustains future competence. We study how routing and training choices maintain useful capabilities in continually updated cascades. In a controlled language-model cascade, upgrading the fallback raises system accuracy from 90.0% to 93.0% while the smaller model's division accuracy falls from 64.4% to 16.2%. Training on the fallback's answers to deferred inputs preserves division accuracy at 65.2%, allowing competence to be maintained even when the fallback remains preferred. Further interventions show that restoring target examples at a fixed training budget brings performance close to the control, while parameter isolation prevents loss from other tasks' updates. An exposure-dependent model of cost-sensitive deferral gives sufficient conditions for a unique competence equilibrium in terms of router smoothness and variation in case-level thresholds, along with exposure requirements and recovery times. Image cascades show a complementary source of competence loss as examples from deferred classes leave the training window. These results support designing routing and training policies around both current performance and retained competence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.