acceptodds
Under review as a conference paper at ICLR 2027

BrowSA: Semantic Aware Sparse Attention for Web Agents with KV cache packing

Abstract

Browser agents re-read tens of thousands of tokens of serialized DOM at every step, yet only a small fraction of that context influences the next action. Sparse attention can skip the rest, but existing methods drag in irrelevant neighbors and/or put substantial pressure on the hardware. We observe that web pages, system prompts, and action histories already carry semantic boundaries through HTML and other structured formats. BrowSA exploits these boundaries twice. It selects among chunks defined by the prompt's semantic boundaries and ranks them with a score calibrated for unequal chunk lengths. A placement stage then packs these chunks into physical KV pages, so the units the scorer selects are the units the kernel fetches. BrowSA reduces decode attention latency by up to 68% over dense attention while matching or exceeding its task success on WebVoyager, GAIA, and Online-Mind2Web.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.