When Does Geometry Add to Recognition? A Fusion Benchmark for Monocular Volume Estimation
Abstract
Monocular object volume estimation is judged by one error averaged over a dataset, mixing two contributions: (i) the recognised category, which already implies much of an object's volume, and (ii) the geometric pipeline that measures it. A benchmark that cannot separate them cannot say which did the work. We introduce OneVol, which factors a monocular estimate into six channels, scores each alone and in combination under rules, combiners, and an oracle on eight datasets, and first asks whether a dataset can support a metric claim. A measurement model gives a bar any estimator must clear, the ground truth's within-class spread, and five validity criteria: five of six volume datasets fail at least one, three on the data alone; and NOCS's 13,371 rows carry the evidence of 46 objects. The same model sizes the next benchmark: about 1/r² targets for a source of within-class correlation r, several hundred to resolve its gain, against the six volume datasets' 9 to 70 targets. The consequence: the median volume of same-class training objects beats the full pipeline in 23 of 24 dataset–backbone cells, and survives a predicted class, unresolved on MADIMA23 and MRGBD. On MetaFood3D, whose classes hold one object each so the prior is only a global median, the best published pipeline beats it. Fused on recognition's terms, the within-class regression keeps the prior and adds only the source's within-class deviation, lowering ECUSTFD's mean log error by 6–11% for every backbone, two of four surviving Holm's correction, and in 25 of 26 cells on the mass datasets, the only ones large enough, confirmed by a pre-registered test on held-out objects and dishes; a second registration, extending the rule to several sources, fails. A weak measurement should be judged by what it adds beyond a strong prior, not by its own accuracy. Code and predictions: https://github.com/nobody-eh/OneVol-review.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.