ROOMER: Learning What to See and When to Stop for Room Re-identification
Abstract
Modern room re-identification pipelines extract increasingly rich global, local, fine-grained, and geometric evidence, but still process every query through a rigid cascade. This uniform execution wastes computation on already-resolved queries and can even overturn correct decisions when later evidence is noisy. We present Roomer, a multi-turn vision-language agent that adaptively allocates test-time evidence, deciding at each turn whether to commit, verify candidates geometrically, or inspect them semantically. To equip the agent with a reliable semantic eye, we introduce Sight-OPSD, which uses training-only regional and geometric teaching views to teach a 4B raw-image Inspector what to see: which cross-view evidence supports or refutes shared room identity. The resulting Inspector requires no privileged inputs at inference and also serves as Roomer’s policy backbone. We train the Router with cost-efficient multi-turn group-relative credit assignment. Beyond trajectory-level Outcome Credit for final correctness, localized Commitment Credit supervises when to stop, while Probe Credit uses evidence displacement to assign credit to semantic inspections that actually improve the evidence state. This localized supervision mitigates an afraid-to-commit collapse in which the policy keeps acquiring evidence instead of answering, while discouraging uninformative use of the expensive semantic Inspector. Roomer improves both recognition and evidence efficiency: Sight-OPSD raises the 4B Inspector’s query-weighted balanced accuracy from 87.80% to 91.12%, exceeding the 397B Qwen3.5 model by 1.46 points, while the full agent reaches 88.79% Recall@1 on AirRoom-Contested, 10.27 points above the fixed AirRoom Full pipeline. On the natural AirRoom distribution, Roomer further reaches 96.46% Recall@1, a 2.19-point improvement over AirRoom Full. Compared with correctness-only routing, localized credit further improves contested Recall@1 by 1.59 points while using 38.8% less tool budget—a Pareto gain showing that Roomer can perform better by looking less. The code will be made publicly available upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.