acceptodds
Under review as a conference paper at ICLR 2027

Auditing Role Collisions in Multimodal Memory Evaluation

Abstract

Multimodal memory evaluation assesses how retrieved evidence supports question answering. However, scores also depend on how this evidence is presented and how responses are generated and scored. We audit a public VisualMem answering path and identify missing explicit role and option bindings, a condition we term a role collision. We introduce role-preserving serialization (RPS) to express these bindings while keeping retrieval outputs fixed. Across four model stacks and all 58 eligible image-choice questions with balanced answer positions, RPS improves accuracy by up to 34.9 percentage points over source-derived serialization, with no gain on Molmo2. To examine how correct bindings affect scores, we separate mapping placement from role wording in a paired 2 × 2 factorial design. On Qwen3.5-9B, raising the output-token limit from 128 to 512 reduces the adjacent-over-global advantages from 25.86 to 5.60 points under neutral wording and from 19.83 to 8.62 under semantic wording; only the neutral reduction passes the expanded Holm correction. By contrast, Gemma and InternVL retain the same answer labels across 64/128/256/512-token runs, while semantic wording lowers InternVL’s global-placement accuracy by 13.79 points. For scoring validation, independent human extraction agrees with the frozen parser on 193/200 archived Qwen outputs. Finally, serialization changes retrieval-method contrasts on 696 questions, while Retrieval-only remains best overall. On 120 MuirBench reference-injection questions, RPS exceeds flat injection by 6.67–27.50 points; these shifts and gains do not pass expanded Holm corrections. Together, these findings show why bindings, output-token limits, and answer-extraction rules must be specified before attributing score differences to retrieval quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.