acceptodds
Under review as a conference paper at ICLR 2027

VERA-ROAD: VISUAL EVIDENCE REASONING AND ADAPTIVE REPAIR FOR ROAD NETWORK EXTRACTION

Abstract

Road-network extraction from very-high-resolution remote sensing imagery is difficult not only because roads vary across scenes, but also because different regions within the same image fail for different reasons. A narrow road may require stronger geometric continuity, a low-contrast segment may benefit from fine- detail enhancement, and a visually ambiguous junction may require a wider spatial context. Most extraction pipelines nevertheless process all regions through a largely fixed computational path. We introduce VERA-Road, a region-wise visual decision framework that replaces this static processing pattern with an observe–act–verify– repair loop. A shared vision-language policy first interprets local, neighboring, and global evidence and issues an action for each region. The proposed Contextual Action Composer (CAC) realizes the action by combining complementary detail, geometry, and cross-band operators. Rather than assuming that the issued action is always sufficient, Outcome-Guided Repair (OGR) measures the resulting prediction state and selectively applies directional-alignment or multi-scale-context repair only to unresolved regions. Finally, Topology-Rewarded Policy Learning (TRPL) optimizes the action policy from downstream graph quality with a group- relative objective, avoiding the need for region-level action annotations. On the SpaceNet road benchmark, the resulting model reaches 86.79 F1, 93.47 precision, 82.51 recall, and 87.42 APLS, while the ablation and sensitivity studies show complementary gains from action composition, selective repair, and topology- aware policy optimization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.