acceptodds
Under review as a conference paper at ICLR 2027

How Deep Is Your Value? Table-Cell Operands Let Go of VLM outputs Earlier Than Answers

Abstract

Vision-language models (VLMs) answer questions about tables directly from document page images, yet final-answer accuracy does not reveal how the visual evidence of a cell is processed inside the model. This matters because a cell’s value can serve two roles: in LOOKUP, it is read out as the answer; in COMPUTE, it is used as an operand with values from other cells, often on other pages. A model could process every cell to the same depth regardless of its role, or release an operand once it has been used. We construct paired LOOKUP and COMPUTE instances that share a table, a target cell, and a pixel edit to that cell. In eight VLMs, we first verify that the cell’s visual tokens carry the value used to construct the answer, and then intervene on these tokens layer by layer. We find that operands let go of the output earlier than answers: replacing the cell’s value redirects the output only up to an earlier layer in COMPUTE than in LOOKUP, in all eight models under a log-probability criterion and in six under full generation. The difference lies in the value, not in the cell’s presence: the depth at which removing the cell stops mattering is set largely by the model and barely moves across question types or page structures. An operand’s value remains linearly decodable at the cell after its influence ends, but later text positions carry the computed result rather than an operand the model can recompute from. Code used for the analysis is publicly available at https://anonymous.4open.science/r/vlm-table-cell-layer-analysis-67D8/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.