ISCon: A Theoretical Framework for Testing Design Rationales of Personalization Models
Abstract
Personalization models are designed around different commitments to time-order invariance, recent evidence, older evidence, and distinct long- and short-term effects. Utility metrics do not establish whether behavior follows those commitments. We propose the *Information-Survival Contract* (), a theoretical framework that takes each model's rationale as given and turns it into a testable behavioral contract. specifies which controlled histories validly test the rationale (*evidence conditions*), how outputs are scored (*measurement*), how scores are combined across evaluation cases (*aggregation rule*), and how the same model's aggregated scores should compare across histories (*acceptable difference*). It distinguishes support, contradiction, and unresolved cases. Its formal properties explain why the same unchanged response can agree with a model designed to aggregate profile evidence with little order sensitivity and contradict one designed to give recent evidence greater influence. We instantiate with , and evaluate models motivated by profile aggregation, recent-state updates, long-range retention, or separate temporal preference components on PENS news recommendation and headline generation, MovieLens-1M movie recommendation, and OpenAI-Reddit personalized summarization. uses two complementary tests under the evidence conditions required by each rationale. One compares the same model's responses to histories containing only long-term (LT), short-term (ST), or episodic (EP) evidence. The other moves, removes, or repeats selected evidence, or introduces distracting or competing evidence. For , whose long-range rationale motivates retaining useful older evidence, LT MRR is , below ST at and EP at , contradicting the tested commitment. A follow-up control finds that late LT placement is MRR worse than deletion. A separate control finds that increasing LT evidence with competing ST evidence fixed raises MRR by . therefore reveals a sharper diagnosis in which older evidence can help this model even though its placement can make that evidence harmful. *The resulting tests identify where a model's own commitments are supported, contradicted, or unresolved without judging which commitment is preferable.*
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.