IntentTwin: Deconfounding Candidate Support and the Oracle-to-Deployment Gap in Offline Web Behavior Prediction
Abstract
Offline web-behavior benchmarks frequently insert the demonstrated action into a finite candidate set and then measure reranking. We show that this protocol can create a candidate-support confound: when the target is outside training support while all distractors are sampled from that support, support membership identifies the gold action without modeling behavior. We formalize this effect with support-identifiability rate (SIR), construct support-balanced candidate sets, and evaluate a target-free dual-path proposer that combines support retrieval with open-vocabulary DOM/LM proposals. On filtered OPeRA (4,282/469/992 train/dev/test steps), the original protocol has SIR on target-OOV rows and a support-only rule reaches OOV Exact@1. Support balancing reduces SIR to ; the strongest conditional reranker obtains overall Exact@1, on IV rows, and on OOV rows. The target-free dual-path system reaches Proposal Recall@16 and end-to-end Exact@1 , compared with oracle conditional Exact@1 . On Mind2Web, the same audit yields SIR and an oracle-to-proposal gap of . A 200-row semantic audit labels of exact-key OOV targets as semantically novel, as locator/string variation, and as ambiguous (Cohen's ). These measurements show why support composition must be controlled alongside candidate accuracy when offline results are used to reason about deployment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.