Compile What Counts: How Question Supervision Shapes Queryable Memory
Abstract
A per-document key–value (KV) code is memory for a frozen reader: optimized once against the document, it is queried in place of the document thereafter. Optimized to store the document, however, such a code need not expose it through the reader's question-answering interface: storage and access are different objectives. What makes the memory queryable is the supervision that bridges them, question–answer pairs that current practice samples from a language model and that guarantee neither per-fact coverage nor question-distribution structure. Question-Grammar Compilation (QGC) instead compiles the supervision: a closed question grammar with three phrasing axes and 27 cells per relation is expanded at exact per-fact exposure, which turns the question distribution into something we can intervene on. A full factorial over axis widths and a path that moves only third-order joint mass at fixed low-order marginals show, on two synthetic platforms, that widening each axis causes gains of up to accuracy with subadditive interactions, while third-order structure and cell-level exactness show no detected response. At matched budget, unconstrained self-study-style generation loses – against the compiled grammar, whereas LLM paraphrases compiled per fact match or exceed it. Axis directions replicate on ClinicalTrials.gov under a prospective gate, where compiled codes reach accuracy against for the zero-shot full document at fewer attention slots than the document. The dominant factor is systematic, fact-grounded supervision with adequate question support; who writes the questions matters only where the grammar's coverage of interrogative forms is thin.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.