Auditing Role Collisions in Multimodal Memory Evaluation
Abstract
Multimodal memory evaluation assesses how retrieved evidence supports question answering. However, scores also depend on how this evidence is presented and how responses are generated and scored. We audit a public VisualMem answering path and identify missing explicit role and option bindings, a condition we term a role collision. We introduce role-preserving serialization (RPS) to express these bindings while keeping retrieval outputs fixed. Across four model stacks and all 58 eligible image-choice questions with balanced answer positions, RPS improves accuracy by up to 34.9 percentage points over source-derived serialization, with no gain on Molmo2. To examine how correct bindings affect scores, we separate mapping placement from role wording in a paired 2 × 2 factorial design. On Qwen3.5-9B, raising the output-token limit from 128 to 512 reduces the adjacent-over-global advantages from 25.86 to 5.60 points under neutral wording and from 19.83 to 8.62 under semantic wording; only the neutral reduction passes the expanded Holm correction. By contrast, Gemma and InternVL retain the same answer labels across 64/128/256/512-token runs, while semantic wording lowers InternVL’s global-placement accuracy by 13.79 points. For scoring validation, independent human extraction agrees with the frozen parser on 193/200 archived Qwen outputs. Finally, serialization changes retrieval-method contrasts on 696 questions, while Retrieval-only remains best overall. On 120 MuirBench reference-injection questions, RPS exceeds flat injection by 6.67–27.50 points; these shifts and gains do not pass expanded Holm corrections. Together, these findings show why bindings, output-token limits, and answer-extraction rules must be specified before attributing score differences to retrieval quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.