ThreadNIAH: Reasoning over Relational Needles in Long Video Haystacks
Abstract
Multimodal large language models can increasingly process hour-long videos, where relevant evidence may be sparse and distributed over time. Effective long-video understanding therefore requires models to locate relevant evidence, link related evidence across the video, and combine multiple pieces of evidence. Building on the controlled setting of Video Needle-in-a-Haystack (NIAH), we introduce ThreadNIAH to systematically evaluate these capabilities. ThreadNIAH organizes five task types into three levels that progressively evaluate locating, linking, and combining evidence: Single-Needle Retrieval (Details), Multi-Needle Threading (Tracking and Comparison), and Multi-Needle Synthesis (Counting and Logic). To instantiate these task types in a controlled and scalable manner, we design 29 task templates that specify the evidence-bearing needles, their relations, and the questions defined over them. Since a correct final answer does not necessarily imply correct intermediate steps, each Main QA is paired with diagnostic Sub-QAs that probe its supporting evidence and intermediate steps. Together, they form a QA chain, the basic evaluation unit of ThreadNIAH. We use Main Acc, Sub Acc, and All Pass Rate (APR) to measure final-task, intermediate-step, and all-question chain accuracy, respectively. The resulting benchmark contains 870 QA chains, comprising 870 Main QAs and 2,580 Sub-QAs. Across ten video MLLMs and five thinking variants, most configurations perform substantially better on single-needle retrieval than on multi-needle tasks. Moreover, thinking does not ensure consistent success across QA chains: even Gemini 3.6 Flash with thinking achieves 67.7% Main Acc but only 49.6% APR. Similar difficulty patterns on real-world long-video questions further support the relevance of ThreadNIAH to real-world long-video understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.