Support Before Sign: Identifying Deferral Gain in LLM Cascades
Abstract
A cascade must decide whether replacing a small language model's answer with a larger model's would improve the graded outcome. Under answer-factored grading, this signed deferral gain factorizes into the probability of answer disagreement and the expected benefit conditional on disagreement. The factors require different supervision. Unlabeled paired outputs point-identify disagreement support but leave the signed margin undetermined. Under canonical exact-match grading and an unrestricted conditional gold-answer law, we characterize the sharp gain interval and show how observationally equivalent outcome laws induce opposite routing orders. Selective outcome labeling identifies the missing margin under conditional independence and positivity. This information structure motivates *Support Before Sign*: learn support from paired outputs without gold labels, learn the margin from labeled disagreements, and route on their product. On a fixed model pair, signed supervision at the main labeled budget improves integrated accuracy over call rate under both the original and equal-weight task mixtures. Its matched-token-cost benefit depends on the mixture, while observation-matched direct estimators perform better at some low-deferral operating points. The results distinguish three questions: what paired outputs identify, what signed labels add, and how the estimator converts that information into a routing policy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.