Was It Never Shown, or Never Seen? Attributing Failure in Observation-Limited 3D Vision-Language Tasks
Abstract
Vision-language systems for volumetric scans, dense 3D scenes, and long videos must construct a budgeted observation before inference. When a prediction fails, aggregate accuracy cannot reveal whether evidence was omitted, the model failed despite adequate evidence, or a useful observation was not selected. We introduce a matched-counterfactual framework for attributing these failures to input-stage (L1), model (L2), and observation-selection (L3) bottlenecks under a declared reference protocol. Cross-domain analysis reveals that observation vulnerability increases sharply under tight budgets and disproportionately affects compact targets. Guided by this diagnosis, we propose a Joint Adaptive Observation Policy that jointly allocates a fixed budget across support, sampling density, representation resolution, and modality-specific mapping. Given a query and a charged low-cost preview, the policy selects a feasible observation. Training combines task reward with dense evidence-adequacy rewards while keeping the task model frozen. It improves low-budget CT grounding, temporal video localisation, and 3D visual grounding over default fixed observations by 26.7, 15.2, and 14.8 points, respectively. Reference-recoverable misses decrease substantially, whereas joint system–reference misses change comparatively little, supporting adaptive observation as an intervention on failures originating before task inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.