acceptodds
Under review as a conference paper at ICLR 2027

Phase Geometry of Reasoning: Predicting and Controlling Evidence Integration in Large Language Models

Abstract

Large language models often produce different probabilistic judgments when equivalent evidence is reordered, reframed, or combined, yet the latent thought dynamics responsible for these changes remain poorly understood. We operationalize a thought as an atomic, belief-changing state transition induced by processing one evidence item, and represent reasoning as a trajectory through a latent state space. We hypothesize that these trajectories follow complex-valued, noncommutative dynamics in which phase-dependent interference determines how successive thoughts combine. To test this hypothesis, we introduce Complex Probabilistic Interventional Reasoning (CPIR), a controlled framework that traces model judgments across three-step reasoning sequences with known priors and likelihood ratios. CPIR intervenes on individual thought updates through evidence masking, polarity reversal, strength modification, order permutation, paraphrasing, and compositionally held-out combinations of these operations. We compare a complex-valued state-space model with normative Bayesian inference, logistic interaction models, nonlinear real-valued recurrent models, and parameter-matched controls. On five frozen synthetic environments with planted complex-coherent dynamics, CPIR identifies the correct compositional family and predicts held-out domains, paraphrases, and longer sequences, reducing joint-holdout logit RMSE from 0.072 to 0.023 relative to the strongest real-valued baseline. When transferred to Qwen3-14B and DeepSeek-V3.2 judgments, the same prespecified procedure discovers target-specific real and heterogeneous predictive regimes, demonstrating selective model assignment across targets. Extending CPIR to sequential LLM evaluation, repeatability and competence controls reveal that the observed cross-order reversal rate (11.1%) is matched by same-order variation (10.0%) for the competence-qualified judge, motivating checklist elicitation that reduces reversals to 0.7%. At equal inference cost, averaging opposite criterion orders improves Brier score from 0.156 to 0.146 over averaging two same-order evaluations, with the gain concentrated in probabilistic calibration. Together, these results show that CPIR can recover structured thought composition, discriminate among compositional regimes, and improve the stability and calibration of sequential LLM evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.