Bridging the Knowledge–Reasoning Gap in Natural Systems: A Benchmark for Inductive Rule and Noise Recovery in Cellular Automata
Abstract
Noise is an inherent part of any natural system. After considering the recent success of transformer-based models in natural system reasoning, this research explores the capability of large language models (LLMs) in addressing the question: given an example of a real-world natural system, can LLMs identify the presence of noise within it? Here, we consider the natural computing model cellular automata (CA) as a proxy of natural systems. Therefore, we formalize this question as the CA identification problem: given noisy evolution dynamics, can LLM recover both the underlying rule and the magnitude of noise? To explore this question, we construct four benchmark datasets (skewECA, αECA, stocECA, and tempECA) totalling 2,876,800 samples across four noise types, along with knowledge-check dataset. This question also captures the reasoning capabilities of LLMs, i.e. reasoning from observation. Evaluating Qwen, Llama, Mistral/Mixtral families (7-72B parameters), alongside GPT 5.1 and Claude Haiku 4.5 under zero-shot, few-shot, and fine-tuned settings (for open source models), we observe near-perfect knowledge accuracy but poor task performance. The rule prediction peaks at 14.60% and noise prediction at 32.50%. This gap is architectural. Next-token prediction does not equip models to extract the global and structural statistics required for this task, including change rates, spatial contiguity, and distributional variance. We therefore design four specialised transformer models (skewM, αM, stocM, and tempM) each aligned with the statistical signature of its corresponding noise process. These models achieve up to % rule exact match and % noise tolerance accuracy, outperforming the best LLM baseline by up to with at most M parameters. These findings emphasize the importance of structure-aware architectures for reasoning over spatiotemporal data, as exemplified by the identification problem. Code and Dataset are available at https://anonymous.4open.science/r/LLM_NOISE-9DC1/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.