acceptodds
Under review as a conference paper at ICLR 2027

The Cost of Reasoning over Frames Instead of Worlds

Abstract

Existing embodied question answering benchmarks evaluate selecting and reading a single frame, analogous to needle-in-a-haystack (NIAH) retrieval. Frontier models excel at this, reflecting their saturation on OpenEQA. We identify an axis of complexity in the scene representation a question requires, defined by how much persistent state must be maintained across frames and how many deductions are needed to integrate observations; existing evaluations occupy its degenerate single-frame case and require neither. We formalize these demands in a compact deductive language whose rules integrate observations across frames into a consistent scene representation, and introduce a novel generative campus environment that poses questions far beyond existing benchmarks along this axis, each certified by a derivation in the language. Frontier Models answer NIAH controls with over 90% accuracy, but fall to 12–41% when evidence must be carried across frames, at derivation depths of at most seven. Results are consistent with models assembling a scene representation ad-hoc within the chain of thought at query time, from the frames themselves, rather than reasoning over one constructed beforehand. Existing benchmarks cannot observe these costs; our axis, language, and environment make them measurable, and motivate evaluations and architectures that treat the representation, rather than the frames, as the object of reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.