The Cost of Reasoning over Frames Instead of Worlds
Abstract
Existing embodied question answering benchmarks evaluate selecting and reading a single frame, analogous to needle-in-a-haystack (NIAH) retrieval. Frontier models excel at this, reflecting their saturation on OpenEQA. We identify an axis of complexity in the scene representation a question requires, defined by how much persistent state must be maintained across frames and how many deductions are needed to integrate observations; existing evaluations occupy its degenerate single-frame case and require neither. We formalize these demands in a compact deductive language whose rules integrate observations across frames into a consistent scene representation, and introduce a novel generative campus environment that poses questions far beyond existing benchmarks along this axis, each certified by a derivation in the language. Frontier Models answer NIAH controls with over 90% accuracy, but fall to 12–41% when evidence must be carried across frames, at derivation depths of at most seven. Results are consistent with models assembling a scene representation ad-hoc within the chain of thought at query time, from the frames themselves, rather than reasoning over one constructed beforehand. Existing benchmarks cannot observe these costs; our axis, language, and environment make them measurable, and motivate evaluations and architectures that treat the representation, rather than the frames, as the object of reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.