Testing Final-State Sufficiency in Multimodal Graph Reasoning
Abstract
A model answering questions about a changing world must track what is true after the updates. When a question depends only on the final state, a model that relies on that state alone should give the same answer distribution for any history reaching it. We call this requirement final-state sufficiency. Accuracy can miss violations when history changes the likelihoods of valid answers without changing the preferred answer. We therefore compare the correct-answer margin, the correct answer's log-likelihood advantage over its strongest valid alternative. Final-state sufficiency requires this margin to be invariant across histories reaching the same state. Our matched graph histories reverse the same two updates while fixing the final graph and answer task; questions about unchanged properties adjust for general order sensitivity. Both Qwen3-VL and Gemma violate the required margin equality in the complete graph on four nodes. A separate held-out study extends the Gemma finding to four larger, non-complete graph families in conditions where the model can answer the graph questions and follow the updates. Correct endpoint answers can therefore conceal history dependence. Dynamic-reasoning evaluations should test invariance across state-equivalent histories.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.