BranchSpan: Continuation-Indexed Replay for Long-Horizon Web Agents
Abstract
Two browser histories can arrive at nearly indistinguishable pages while leaving different actions executable because an earlier tab, filter, modal choice, or selection changed the active branch. This branch aliasing makes replay an addressing problem: chronological context preserves history and semantic retrieval finds relevant-looking states, but neither address determines whether recalled actions still fit the branch on screen. We introduce BranchSpan, which seals decision-coherent trajectory segments and indexes each endpoint by latent relevance, canonical reachable actions, and a history-conditioned locality-sensitive branch code; the selected segment enters an otherwise unchanged agent through a deterministic five-field payload. On the complete 812-task WebArena and 910-task VisualWebArena inventories, BranchSpan improves GPT-4o success over 128k Long-Context by +5.3 [3.6,7.0] and +4.4 [2.8,6.0] points, rising to +7.2 [4.6,9.8] and +5.9 [3.6,8.2] on long-horizon tasks. A validation-tuned Long-Context plus semantic-retrieval hybrid leaves BranchSpan ahead by +2.6 [1.1,4.1] and +2.3 [0.9,3.7] points overall. In a method-blind audit of method-native replay, annotators find valid continuations in 82.1% (197/240) and 76.7% (184/240) of sampled WebArena and VisualWebArena states; on the shared WebArena segment inventory, replacing continuation scoring with cosine-only retrieval reduces success by 1.9 points overall and 2.8 on long tasks, and reduces valid replay by 5.8 points. The frozen interface preserves the Long-Context hybrid BranchSpan ordering with Qwen2.5-VL-72B, while BranchSpan uses 0.21 the Raw Frame replay payload with 40.3 ms encoder cost. For long-horizon agents, a useful memory address must recover not only relevant history, but history whose actions remain executable now.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.