Reasoning in Life: Can Language Agents Master Evolving Worlds?
Abstract
Reasoning is the core capability for large language models (LLMs) to solve complex tasks such as planning and control. Such reasoning tasks can be categorized into three forms: deduction, which predicts the future of a system from its rule and current state; induction, which recovers the hidden rule from observed transitions; and abduction, which reconstructs the antecedent that best explains an observed outcome. Existing benchmarks, however, either consider a single form of reasoning or lack the formal grounding across the three forms, leaving no principled way to measure the divergence of language agents over different reasoning forms. To narrow the gap, we introduce LifeWorld, a benchmark built on life-like cellular automata with rules generalizing the famous Conway's Game of Life (GoL), where the evolving step can be characterized by the rule , a configuration , and its -step successor . We curate the dataset for three forms of reasoning: i) deduction () is forward simulation with a unique answer, computable in polynomial time yet Turing complete in expressive power; ii) induction () is rule identification from observed transitions, corresponding to program induction from input-output examples; and iii) abduction () is predecessor reconstruction under irreversible dynamics, which is NP-complete at bounded horizon. We further propose LifeAgent, a recursive reasoning harness that operates through two forms of recursion: i) space-time recursion, which breaks a query into sub-grids and sub-horizons, and ii) logical recursion, which solves induction and abduction by a propose-then-verify loop in which candidates are checked through deduction. Extensive experiments across frontier LLMs demonstrate the effectiveness of both recursions in LifeAgent and provide a comprehensive analysis of performance across the three reasoning forms, characterizing the behaviors of frontier LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.