TenonBench: Evaluating Spatiotemporal Capabilities with Mortise-and-Tenon Assemblies
Abstract
Evaluating whether vision-language models (VLMs) can reason about feasible motions and operation order is essential for their use in embodied planning. However, existing benchmarks provide limited coverage of changing geometric constraints, do not fully distinguish necessary from optional ordering, and often assess perception and planning separately. To address these gaps, we introduce TenonBench, a spatiotemporal reasoning benchmark that combines interlocking assembly tasks, multi-solution evaluation, and paired perception–reasoning analysis. We construct 178 parametric mortise-and-tenon assemblies with 1,200 parts and use a geometric solver to provide ground truth for 4,065 questions spanning part identification, structural understanding, and reasoning about how movement direction and preceding removals determine motion feasibility. Our partial-order evaluation distinguishes necessary dependencies from optional ordering that varies across valid solutions, while complete sequences are verified by replaying each removal against the remaining assembly. Within multi-turn shared conversations, we further pair reasoning questions with perception probes about the same parts and jointly analyze their responses to examine whether reasoning errors persist when the corresponding perception questions are answered correctly. Evaluation of 23 VLMs yields a best overall accuracy of 77.2% and a mean of 53.7%, highlighting TenonBench as a challenging testbed for geometry-grounded spatiotemporal reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.