acceptodds
Under review as a conference paper at ICLR 2027

The Unit of Memory Supervision Is the Set

Abstract

Memory selectors for language agents score each candidate item on its own, by relevance, salience, or recency. But a memory helps only through the set it forms, and a set’s value is rarely the sum of its items, so which items to keep is a question about the set. We show that item-level supervision cannot in general identify the optimal set. Two value functions can have identical add-one and delete-one effects for every fact yet have disjoint optimal budgeted sets, so no rule restricted to those add-one and delete-one signals can avoid set regret. On a sealed split of 250 MuSiQue questions with exact subset enumeration, a development-selected rank-sum of these singleton signals leaves 0.18 answer-F1 regret against the optimal set, while a set-derived item score leaves 0.006. Changing only the training label from salience to gold supporting sets, at matched architecture, data, and budget, raises exact match from 0.107 to 0.345. On 500 fresh MuSiQue questions with retrieved candidates, supporting-set supervision beats both add-one item supervision and dense retrieval. The supervision must observe value at the set level, but deployment need not: set-level value can be distilled into per-item scores, preserving item-wise selection at test time. When downstream value is interactive, the unit of supervision is the set.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.