acceptodds
Under review as a conference paper at ICLR 2027

PixelRAG: Retrieval and Reading in Pixel Space over Millions of Web Screenshots

Abstract

Augmenting large language models (LLMs) with retrieved web text has become a dominant paradigm, yet the web is not natively textual: existing systems depend on complex parsing pipelines that linearize HTML and discard layout, visual structure, and formatting. We introduce PixelRAG, a new retrieval-augmented method that represents websites in their native visual form and performs retrieval and reading entirely in pixel space, enabling an end-to-end architecture that eliminates text abstraction. PixelRAG is, to our knowledge, the first pipeline to operate over a full Wikipedia corpus in this form, scaling to a datastore of 30 million screenshot images with an efficient visual retrieval index. Built on an existing visual embedding model (i.e., Qwen3-VL-Embedding), PixelRAG further fine-tunes this model on screenshot data with carefully curated contrastive training data. Retrieved screenshots are then fed directly as pixel inputs to a VLM, without intermediate text conversion. PixelRAG consistently outperforms both no-retrieval and text-based RAG baselines, improving accuracy by up to 18.1% over text-based baselines, most surprisingly on widely studied text-centric tasks such as NQ and SimpleQA. It also achieves strong gains on multimodal open-domain QA (e.g., MMSearch), benchmarks over noisy news corpora (e.g., LiveVQA), and agentic benchmarks (e.g., MoNaCo). Finally, pixel representations enable an efficiency lever for RAG through image compression, effectively reducing reader input tokens by up to 3× while maintaining accuracy. Our results question the necessity of text representations in web retrieval, suggesting that web RAG can operate in the web's native visual form while improving accuracy and reducing reader input tokens.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.