acceptodds
Under review as a conference paper at ICLR 2027

Local Tests, Long Horizons: What Memory-Compaction Tests Identify

Abstract

Short reset tests can be reused to assess repeated memory updates, but a precise local test need not identify long-horizon risk. We study which output statistics a test must retain, which uncertainty a computable bound actually covers, and when the resulting test is affordable. Complete marginal tables can admit nearly unit stationary risk ambiguity. A support-aware continuation criterion characterizes exact identification for all initial states, while quantitative examples show why small first-horizon ambiguity is not a deployment-horizon guarantee. Across 330 geometries, certified frozen-kernel brackets reveal discrepancies that a two-step diagnostic misses. We then select among candidate output features using simultaneously calibrated observations, without supplying the generator's influential feature. Selection improves some matched-budget certification decisions but loses at the smallest budget because of its statistical correction. Established random-set and robust dynamic programming (DP) methods provide propagation; we make their information, temporal, and sampling relaxations explicit. The supported operating regime is cheap resettable software with declared small-block structure, not unrestricted language-model compaction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.