BINDMEM: Executable Memory for Long-Horizon Reasoning
Abstract
Long-horizon agents must interpret and combine information across past interactions, yet access to supporting evidence does not ensure correct reasoning. Even a successfully computed value may fail to satisfy the question's required scope or precision. We introduce BINDMEM, a neuro-symbolic memory interface that preserves source-local interpretations (WRITE), executes query-dependent programs (READ), and checks answer requirements before disclosing derived values (DISCLOSE). WRITE materializes source-local interpretations as provenance-linked typed records whenever their inputs are available from the source context. READ compiles each question into a typed program and an answer contract, then executes it over the retrieved records while tracking supporting evidence and unresolved conditions. DISCLOSE supplies a computed value only when it meets the requested precision and scope; otherwise it supplies the selected evidence alone. The final reader retains the original retrieved text throughout. Across six benchmark variants with two instruction-tuned readers, BINDMEM attains the highest overall score in all 12 comparisons against seven memory baselines, and it also improves a frontier reader. Fixed-retrieval controls isolate the contributions of stored interpretations, execution-selected evidence, and selective disclosure, and with identical stored records BINDMEM outperforms the evaluated Python pipelines for both readers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.