Retrieval-Augmented Transparent Object Perception for Autonomous Laboratories
Abstract
Transparent objects are ubiquitous in autonomous laboratories, yet their geometry remains difficult to perceive due to refraction, reflection, and unreliable visual cues. Existing methods largely infer transparent geometry from observations alone, overlooking a distinctive property of laboratory environments: they are semi-open worlds, where many instruments are standardized products with known or retrievable 3D models. We introduce Retrieval-Augmented Transparent Depth Estimation (RATDE), a framework that grounds visual depth recovery with external object geometry. Given a transparent laboratory object, our approach retrieves its corresponding CAD model, encodes the CAD geometry as point-cloud features, and establishes image-to-CAD correspondence to obtain pose-aware local geometric tokens. These geometric priors condition a transparent depth model, enabling it to jointly reason over ambiguous visual evidence and known 3D structure. Importantly, the framework uses pose supervision during training while requiring no ground-truth pose at inference time. To study this problem at scale, we construct a new laboratory transparent-object benchmark, LabSense-Glass, containing 231 laboratory instruments across 222 laboratory scenes, explicitly associating visual observations with retrievable 3D assets. We evaluate depth recovery together with CAD retrieval and 6D geometric alignment and further study generalization under unseen objects, missing CAD models, and imperfect retrieval. Our results demonstrate that exploiting retrievable geometry provides a strong prior for transparent perception, suggesting a broader principle for robotic laboratories: when geometry already exists in the world, perception should retrieve and ground it rather than infer it entirely from scratch.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.