acceptodds
Under review as a conference paper at ICLR 2027

Was It Never Shown, or Never Seen? Attributing Failure in Observation-Limited 3D Vision-Language Tasks

Abstract

Vision-language systems for volumetric scans, dense 3D scenes, and long videos must construct a budgeted observation before inference. When a prediction fails, aggregate accuracy cannot reveal whether evidence was omitted, the model failed despite adequate evidence, or a useful observation was not selected. We introduce a matched-counterfactual framework for attributing these failures to input-stage (L1), model (L2), and observation-selection (L3) bottlenecks under a declared reference protocol. Cross-domain analysis reveals that observation vulnerability increases sharply under tight budgets and disproportionately affects compact targets. Guided by this diagnosis, we propose a Joint Adaptive Observation Policy that jointly allocates a fixed budget across support, sampling density, representation resolution, and modality-specific mapping. Given a query and a charged low-cost preview, the policy selects a feasible observation. Training combines task reward with dense evidence-adequacy rewards while keeping the task model frozen. It improves low-budget CT grounding, temporal video localisation, and 3D visual grounding over default fixed observations by 26.7, 15.2, and 14.8 points, respectively. Reference-recoverable misses decrease substantially, whereas joint system–reference misses change comparatively little, supporting adaptive observation as an intervention on failures originating before task inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.