acceptodds
Under review as a conference paper at ICLR 2027

Language Models Track State but Underuse Observation Reliability

Abstract

Reliable state estimation requires adjusting each update to the quality of new evidence. We study this adjustment in tasks with exact Bayesian references. A five-model survey reveals strong posterior-mean tracking but weak adaptation to sensor reliability. We introduce a paired audit that holds the displayed history fixed and perturbs only the current reading, measuring its influence without regressing on previous model outputs. Under greedy decoding, three checkpoints from two families realize only 5-8% of the prescribed gain change in a matched perturbation test as sensor variance increases sixteenfold. Focused experiments reveal that the same checkpoint can produce accurate updates: supplying the predicted state improves posterior accuracy, and a reasoning-enabled protocol sharply reduces both local-gain and posterior-mean errors on held-out trajectories with unchanged prompts. We translate this diagnosis into the Perception-Dynamics Decoupled Agent (PDDA), combining language-based report extraction with explicit Kalman filtering. Complete-system and matched-extraction comparisons improve monitoring accuracy without generated reasoning at the numerical update stage. Reliability-conditioned updating thus provides a diagnostic beyond trajectory agreement and a concrete target for improving language-based estimation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.