acceptodds
Under review as a conference paper at ICLR 2027

CogExtract: Test-Time Search with Execution-Grounded Verification for Web Information Extraction on Dynamic Pages

Abstract

Web information extraction systems must generate reusable rules that remain reliable across pages with locally varying structures. Existing approaches typically generate a single rule and reuse it without external verification, making errors difficult to detect and repair. We introduce CogExtract, a test-time search framework that couples three mechanisms: Diversity-Conditioned Generation constructs structurally distinct XPath candidates; Execution-Grounded Verification ranks them using deterministic execution signals from a seed page and a sibling verification page; and Failure-Conditioned Reflection uses the resulting failure evidence to guide targeted regeneration when no candidate passes verification. We evaluate CogExtract on LiveWeb-IE, covering 14 websites, 312 page groups, and 50,524 URLs, with 11 vision-language model backbones. CogExtract achieves the best F1 for every backbone and improves over VGS by 14.5 points on average, while increasing model calls by only 13%; GPT-5 reaches the highest overall F1 of 76.6. Mechanism-level analyses show that execution-grounded scoring outperforms LLM self-evaluation in candidate discrimination (AUC 0.840 vs. 0.640; best-selection P@1 96.2% vs. 77.9%). Within the audited a/c failure subset, structured execution-grounded diagnosis also achieves 91.0% accuracy, compared with 59.0% for LLM judgment. These results demonstrate that external execution evidence provides an effective basis for both rule selection and targeted repair on dynamic web pages.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.