acceptodds
Under review as a conference paper at ICLR 2027

RxnHaystack: Beyond Needles in Long-Context Structured Scientific Data

Abstract

Large language models (LLMs) deployed in AI-for-science increasingly confront structured datasets that far exceed their context windows. Existing long-context benchmarks, however, rarely distinguish failures of dataset-scale access, systematic execution over structured objects, and domain-specific reasoning. We introduce RxnHaystack, a diagnostic benchmark over 122,456 cleaned USPTO reactions. Its 100 questions span exact retrieval, deterministic aggregation, reaction-level inference, and relational reasoning over synthesis graphs. We evaluate plain prompting, CodeAct, and Recursive Language Models (RLMs) across 21,500 question-level trajectories. Plain-prompting performance declines in every tier as the supplied context grows from 100 to 500 reactions. CodeAct remains stable on Tiers 1–2 but declines on the harder Tiers 3–4 as contexts extend to 1,000 reactions. Both interfaces are bounded by the model context window, preventing them from receiving very large corpora in full. Recently proposed RLMs address this access constraint through recursive processing beyond the prompt window. Their full-corpus setting exposes agents to all 122,456 reactions, enabling more demanding corpus-wide questions and revealing which challenges remain once the complete dataset is accessible. Across Qwen3.5-397B-A17B, DeepSeek V4 Flash, Gemini 3.7 Flash, GPT-5 mini, and Claude Haiku 4.5, mechanically specified graph operations average 0.81 macro-F1, while mechanism classification averages 0.34 and route and multi-constraint tasks average 0.08. In controlled tests, chemistry-intensive performance falls as more reactions must be searched, even when every corpus contains one correct reaction. Providing the correct chemistry rule, but not the answers, improves classification, yet models still fail to apply it consistently across the entire corpus. Together with deterministic ground truth and inspectable traces, these controls turn aggregate performance into a diagnosis of why an agent fails, providing a reproducible testbed for agents that must operate systematically over large structured scientific data.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.