acceptodds
Under review as a conference paper at ICLR 2027

AuTO-Pose: Autoregressive Transparent Object 6D Pose Estimation with Ray-Background Context

Abstract

We present AuTO-Pose, a novel framework for transparent object 6D pose estimation from RGB videos that causally decodes each target frame window conditioned on previously processed frames. Existing state-of-the-art methods for transparent object pose estimation either adapt opaque object pose estimators or learn cues from isolated transparent observations, largely overlooking temporal context and remaining vulnerable to background-dependent appearance variations. In contrast, AuTO-Pose predicts poses causally from streaming visual observations while accumulating object-centric evidence from previously processed frames. By constructing context from ray-background interaction cues across views, AuTO-Pose exploits how transparent objects refract, reflect, attenuate, and rearrange background evidence over time. This allows the model to resolve ambiguous single-frame cues and generalize better to challenging scenarios with changing backgrounds, weak boundaries, occlusion, and long sequences. We further adopt symmetry-aware supervision to avoid inconsistent training signals for transparent objects with indistinguishable orientations. Extensive experiments on TRansPose, ClearPose, and our self-collected TOS-6D benchmarks show that AuTO-Pose consistently improves transparent object pose accuracy and achieves state-of-the-art performance across benchmark accuracy, sequence generalization, and scenario-level robustness protocols. More details can be found on our [project page](https://anonymous.auto-pose.pages.dev).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.