ChainQABench: A Benchmark for Evidence-Grounded Question Answering over Blockchain Data
Abstract
Public blockchains make every transaction observable, yet even simple economic questions, such as how much an address gained, require combining many raw records. Doing so involves decoding the records into protocol actions, grouping them into analytical units, and judging what the result can establish. Large language models (LLMs) find these steps hard to perform directly on raw records, where values are hex-encoded and a single action spans several logs and traces. To address this, we introduce ChainQABench, to our knowledge the first evidence-grounded question-answering benchmark for blockchain data. Its 160 expert-curated questions cover asset flows, exchange attribution, value extraction through transaction ordering (MEV), and wash-trading patterns on decentralized exchanges. Five domain experts wrote rulebooks fixing admissible evidence and when a claim must remain indeterminate; an independent verifier reproduces every reference from its raw records. Over 16M transactions and 105M traces, we build an evidence hierarchy that organizes raw records into protocol actions, analytical units, and attribution context while preserving provenance links. We evaluate eleven LLMs as evidence moves from raw records to protocol actions, analytical units, and retrieved evidence. Even the strongest LLM, Gemini-3.8-Flash, achieves a maximum question accuracy of only 53.1% across these evidence conditions. Decoding raw records into protocol actions yields the most reliable gains, exposing typed quantities that LLMs otherwise struggle to reconstruct. We release the benchmark, evidence hierarchy, and independent verifier to support research on how evidence representations shape on-chain reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.