acceptodds
Under review as a conference paper at ICLR 2027

Machine Causality Stratification: Derived Authority, Reversible Causal Ledgers, and Coverage-Gated Capability

Abstract

Memory systems for agents increasingly store causal claims, yet the causal lines surveyed in Section 2 report capability as a single level rather than a graded, evidence-derived one. That framing cannot answer what matters when a system’s own memory encodes causal claims: how deep is the claim, why should it be believed, can its authority be revoked, and can one high-level edge make the whole system look causally mature. We present CKB, a causal-memory substrate with a machine-governed causal-authority lifecycle: a graded CDL (Causal Depth Level, associative to interventional), a predicate-derived and reversible CAL (Causal Authority Level, a strict accumulation over evidence predicates), a per-edge capability coordinate CCC, and a coverage-gated envelope CCE that stops one high-depth edge from inflating the reported depth. Authority is captured in an append-only, reversible witness DAG whose transitions are receipted and replayable. We separate what kind of causal claim an edge expresses (CDL) from how much authority the system may assign to it (CAL). An implemented Rust substrate and a cross-language Python reference oracle execute the calculus at the substrate and conformance layer; we report structural certificates, exact indexed recall, a registered counterfactual edge audit, and downstream measurements including a null result for incremental explicit-structure presentation. Downstream, a preregistered paired test (N = 80, a LongMemEval-S public-development subset) finds that external memory helps substantially (+32.5 pp over no memory, p = 3.0e-8), but that presenting the same evidence as typed causal/temporal structure yields no detectable gain over a flat presentation (+1.25 pp, p = 1.0); the null persists under a matched-token placebo arm. On the full 500-question public development set, the frozen official gpt-4o-2024-08-06 evaluator scores a frozen candidate at 493/500; with no baseline and no recorded answer generator, we report this as an observability result, not a comparative or generalization claim (Section 8).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.