acceptodds
Under review as a conference paper at ICLR 2027

ContractShift: Fixed-Trajectory Scoring Reveals Memory-System Selection Regret

Abstract

Memory-system selection depends on the state obligations encoded by its evaluator. ContractShift makes this dependency auditable by applying registered retrieval-heavy and mutable-contract scorers to identical multi-session trajectories while systems, prompts, outputs, task IDs, and budgets remain fixed. We formalize the family- and seed-level estimators needed to compare scorers without changing the realized workload, and specify six observable state-operation contracts: store, retrieve, revise, deprecate, erase, and prove-origin. The instrument fixes task support before aggregation, treats task ID as the paired inference unit, and distinguishes an empirical scorer comparison from native current-state evaluations. We instantiate the protocol on 1,080 multi-session tasks and seven architectures, report native Memora and MemoryAgentBench outcomes with their stated raw counts, validate the harness against independent human judgments, and use a fixed and an adaptive residue audit to characterize bounded recovery. We do not report a fixed-trajectory score, near-best verdict, gap, confidence interval, or selection-regret estimate without the task-level scorer ledger needed to reproduce it. ContractShift therefore provides a contract-sensitive evaluation design for determining when answer-level scoring can mask revision, invalidation, erasure-residue, and provenance obligations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.