acceptodds
Under review as a conference paper at ICLR 2027

ChronoCheck: Scale Improves Per-Claim Faithfulness but Shifts Time-Series Rationales Toward the Claims Models Get Wrong

Abstract

Larger language models are more faithful on a fixed mix of claim kinds when they reason about time series, and they spend that gain on the kinds of claim they get wrong most. We measure this with ChronoCheck, a training-free oversight layer that decomposes a rationale into typed atomic claims (trend, extremum, periodicity, change-point, comparison, anomaly) and compiles each into a deterministic statistical check over the raw samples, so a rationale is refuted claim by claim. The checkers are conservative on the contradiction side: on a fresh suite of certified-true synthetic claims the contradiction criterion produced 0 false positives, and what they certify is consistency with the extracted claims rather than the truth of the rationale. Across three frozen scales chain-of-thought accuracy rises from 55.5% at 7B to 63.6% at 32B while the pooled contradicted-claim rate does not fall; held at the 7B mix of claim classes it falls from 36.5% to 31.5%, because claims naming a value or a location, the classes refuted most, take up a growing share of the rationale. Repair does not close the gap: no repair delta separates from zero at any Qwen2.5 scale (one ungated arm does on Llama-3.1-8B), and a large share of refuted claims disappear from the rewrite rather than being restated. Certify-abstain is the only arm to reach zero contradicted claims, at low coverage, with certified-subset accuracy gaps whose bootstrap intervals all cross zero. The checker measures time-series rationales; it does not repair them. We release it with the claim suite.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.