Recursively Updated Summaries Can Amplify a Remembered Betrayal’s Effect on a Qwen3-8B Agent’s Stated Willingness
Abstract
Long-horizon language-model agents compress their interaction history into summaries. We ask whether a single remembered relational event changes an agent’s later willingness to rely on the offending party, and whether two common summarization procedures preserve or distort that change. In a scripted three-agent simulation on Qwen3-8B, with a scenario family (n = 30 lexical variants of one script) as the unit of inference and a single-candidate digit-logit rating as the instrument, a deliberate betrayal on day 4 lowers Alice’s willingness toward Bob by 0.46 on a 7-point scale; the gap between Bob and Charlie, the recipient of the disclosure, falls by 0.136 relative to neutral control (95% CI [−0.158, −0.117]; 30/30 families). An initial comparison suggested that both a recursively updated summary and a summary regenerated from the archive amplified this effect. A content audit revealed an instruction asymmetry between the summarizer prompts; in a preregistered replication with matched prompts the regenerated-summary result disappears on the checkpoint average, while a smaller recursive-summary amplification remains (−0.055, CI [−0.075, −0.037], 27/30 families negative; the originally secondary contrast). A preregistered sentence-swap study shows that the verbatim wording of the retained event sentence does not by itself produce this difference; the responsible feature of the surrounding text is not isolated, so we report a difference between recursively updated and regenerated memory states, not an isolated effect of recursion. Minimal edits to the memory text show the effect is bound to whichever name occupies the offender role and depends on the confidentiality-request clause; a per-candidate decomposition shows that compact “judgment” sentences widen the gap mainly by raising willingness toward the comparator, a caution for gap-based instruments. On the original corpus, explicit recall of the event stays intact. The instrument could not be validated on other model families in our calibration protocol, so all findings are limited to one model and lexical variants of one scenario.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.