Learning Affective Memory for Behavioral Alignment in Long-Horizon Agents
Abstract
Long-horizon agents must preserve the behavioral implications of earlier user concerns as their interaction history is compressed. Affective cues such as urgency, frustration, shame, anxiety, overconfidence, or distrust can indicate what is at stake without changing the task facts. Losing their relevant implications can produce grounded but behaviorally misaligned actions; retaining them indiscriminately can cause affect lock-in, leakage into neutral steps, or over-compliance with unsafe pressure. We study this problem as Memory as Affective Control (MAC). Our method is a trainable additive module that attaches to a semantic memory backbone and constructs compact, updateable affective memory encoding the action implications of user concerns: stance, behavior-control knobs, step-level gates, update/masking rules, and preferred or avoided paths. It preserves task-relevant affect, dampens stale or excessive affect, masks distractors, and retargets behavior toward alternative valid policies. We construct ambiguity-calibrated long-horizon coding and deep-research benchmarks where facts support grounding but do not determine the desired policy, with supervised and preference data for learning affective memory construction and execution. Across multiple memory backbones, compact affective memories improve behavioral steering over assigning the same total token budget to additional semantic memory, with high judge-rated task correctness and grounding. On Terminal-Bench 2.1 and BrowseComp with matched synthetic affective context, we compare against an equal-budget structured control summary with the same gate and update opportunities. Method-blind behavior diagnostics, human preferences, and compaction analyses support the updateable residual; native accuracy shows no observed mean trade-off in the tested runs. The results support memory as a means of preserving task-specific behavioral alignment across long-horizon execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.