acceptodds
Under review as a conference paper at ICLR 2027

SIGMA: Label-Efficient Training of Language Models to Follow Changing Structured Guidance

Abstract

Language models offer a flexible interface for task-specific decisions, yet accuracy under fixed instructions does not establish reliable behavior when those instructions change. We investigate how limited task annotations can teach models to follow updates that change the required decision while remaining stable under changes that preserve meaning. We introduce SIGMA, which pairs each labeled input with two versions of structured guidance and supervises both correct answers and the expected relationship between their predictions. By aligning the relative scores of all candidate outputs across guidance versions, SIGMA reuses existing annotations to learn how decisions should respond to guidance edits. On BANKING77, CLINC150, and HWU64 with Qwen2.5-7B-Instruct, SIGMA trained on four examples per intent exceeds the mean accuracy of supervised low-rank fine-tuning using sixteen examples per intent under unseen guidance combinations on every dataset, using 75% fewer training annotations with the same development set. The results support supervision across guidance versions as a way to improve annotation efficiency and handle unseen combinations of updates without further training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.