TRUSTCONS: EXACTLY VERIFYING CONSISTENCY, COMPLETENESS , AND CONDITIONED RE-ENUMERATION IN LANGUAGE MODELS
Abstract
Reliable constraint reasoning requires more than a correct satisfiability verdict: a model should recover the target set, preserve answers under meaning-preserving changes, and — when told that a candidate set contains exactly one deterministic error — return the exact target set. We introduce TRUSTCONS, an exhaustive- oracle diagnostic with 48 boundary-selected Boolean instances, five views, and three tasks, plus a matched extension that holds visible formula dimensions fixed while raising the exact solution count from 1 to 32. An exploratory run exposed a completion-budget confound; we froze it, reran complete matrices under larger limits, treated cap-hit unparseable outputs as missing, and reverified every solution set with Z3 5.0.0. Because v2 reissued complete matrices rather than selected cells, 7,344 cells carry both a v1 and a v2 record under one configura- tion, instance, view, task, and prompt, so the correction is measured rather than asserted: the larger caps moved 3,184 censored cells into the observed popula- tion, 2,839 of them to strict correctness, a paired change of 38.4pp (95% base- cluster interval [34.5, 42.1]pp). Unconditional worst/best-case macro ranges are 45.7–68.2%, 50.2–57.8%, and 25.7–64.2% for decision, low-cardinality exact-set recovery, and the candidate-conditioned third task; the pooled point estimates in- side them are record-level averages that include the below-floor cells, and equal- weighting configurations yields different values. Among 192 fully observed de- cision/enumeration pairs from 4 configurations, 5.2% (95% base-cluster interval 2.1–8.9%) are decision-correct but enumeration-wrong. The third task returns the model’s own enumeration set on 89.8% of strictly parsed pairs and, on those 147 parsed pairs, is correct about equally often (7 against 7, exact McNemar p = 1.00; unparseable answers are missing, not incorrect), so it is scored as conditioned re- enumeration, not as evidence of detection, localisation, or repair. TrustCons is a small synthetic diagnostic, not a safety certificate; its contribution is protocol- aware joint measurement and a descriptive stress test of exact-set scaling under matched visible dimensions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.