acceptodds
Under review as a conference paper at ICLR 2027

Beyond Recognition: Benchmarking and Advancing Fine-Grained 3D Object Understanding

Abstract

Object-centric understanding is fundamental to how intelligent systems perceive, reason about, and interact with the 3D world. Yet 3D vision-language models (VLMs) have largely been evaluated on category recognition, captioning, and broad question answering, which reveals limited insight about whether they understand the fine-grained properties of individual objects. General-purpose VLMs offer a promising alternative by reasoning over multi-view renderings, but it remains unclear whether they can answer fine-grained, object-specific questions reliably, and whether that reliability persists when an asset’s appearance or orientation changes. We introduce ObjectArena3D, a benchmark of 16,319 questions over 12,431 assets spanning 20 tasks across semantic attributes, camera awareness, part structure, and surface properties. Beyond standard evaluation, it tests robustness by randomizing object orientations and replacing textures with random solid colors. Our evaluation reveals weak viewpoint and orientation reasoning, with category recognition further degraded under randomized orientations. To narrow these gaps, we propose Objective3D, an framework that fine-tunes a pretrained VLM using view augmentation and intra-view attention across aligned rendering modalities. Compared to baseline Qwen3-VL-8B, Objective3D raises the average benchmark score from 42.95 to 70.26, surpassing the strongest proprietary baseline GPT-5.6-Sol by 4.34 points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.