Mind the Supersession Gap: Diagnosing Event-to-Preference Inference Failures in Language Models
Abstract
Alignment between a user and a personalized assistant is not a static achievement. Circumstances evolve over time, and users typically communicate these shifts through life events rather than explicit declarations. Whether a preference revision successfully registers with the assistant is therefore a property of the coupled interaction system rather than a capability of the model in isolation. We introduce a novel diagnostic framework that rigorously isolates implicit event-to-preference inference from standard memory retrieval. Across an evaluation of nine open-weights models and four API models, we find that assistants honor explicitly declared revisions with near-perfect accuracy but rarely recognize revisions implied by events. Performance falls significantly below chance at 2.2 to 22.1 percent. These implicit signals are nevertheless highly legible to humans. Annotators resolve the same updates at 100 percent accuracy on items where they endorse the premise. We demonstrate that this deficit is neither a long-context retrieval failure nor a bottleneck solvable by increased test-time computation. Blind-scored reasoning traces reveal a dissociation between stated reasoning and final choice. Models that derive an event's consequence are no more likely to act on it than those that do not. Mechanistic analysis using the Jacobian lens shows that the event-implied value emerges early but loses influence by the answer. Contrastive steering and activation patching can recover implicit updates in some settings, but no intervention tested provides a general repair because increasing uptake of the inferred preference can also weaken adherence to explicitly stated preferences. The strongest predictor we identify is the weight the interaction places on the initial statement, which is determined by the framing. Recasting the exchange as an ordinary conversation improves accuracy by 52 points, while inducing a mirror vulnerability where assistants improperly override explicit preferences. A deployed system can thus be overly deferential to initial statements and insufficiently deferential to recent ones. We have provided the datasets, model output, and code.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.