acceptodds
Under review as a conference paper at ICLR 2027

Learning What Not to Say: Conditional Semantic Control in Language Models

Abstract

Language models can follow instructions about what to say, yet remain unreliable at controlling what not to say. We study semantic omission: completing a writing task while keeping an instructed semantic dimension absent from the generated text. Unlike surface-level constraints, semantic omission requires preventing the target semantics from reappearing through paraphrases, implicit references, and rhetorical structures, while preserving the utility of the remaining content. We construct a benchmark of 400 writing tasks across six domains and four writing settings, and systematically evaluate semantic omission across multiple model families. Our diagnosis reveals a behavioral gap between explicit boundary identification and artifact-level omission: models may correctly identify what should be excluded, yet still reintroduce the excluded semantics during generation. Thus, successful boundary identification alone is insufficient for reliable semantic omission. Motivated by this behavioral dissociation, we formulate the repair problem as conditional semantic control and propose SelectOmit, a post-training method that teaches models to selectively control whether a semantic dimension is expressed according to the current instruction. It jointly encourages suppression of the excluded semantics, preservation of the remaining task content, and recovery of the same semantics when explicitly requested. Experiments across multiple model families show that SelectOmit substantially improves semantic omission while preserving task utility, and enables reliable instruction-conditioned inclusion and exclusion of the same semantic content.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.