acceptodds
Under review as a conference paper at ICLR 2027

Rehydration: Fixed-Semantics Counterfactual Replay for Mechanistic LLM-Agent Social Simulation

Abstract

Counterfactual experiments with large language model (LLM) agents are diffi- cult to interpret when an intervention changes the platform state, thereby altering subsequent prompts and generated language. Standard replay comparisons can therefore mix structural policy response with post-intervention semantic adaptation, even when branches share the same decoding seed. We formalize behavioral and platform interventions through the Behavior–Platform Decomposed Mechanistic Transition Framework (BDMTF) and introduce Rehydration, a paired coun- terfactual replay protocol that generates and seals semantic opportunities before policy assignment and reuses them across intervention branches. This construction defines a fixed-semantics structural effect that complements the conventional adaptive-semantic effect. In an analytic setting with known ground truth, Re- hydration recovers the fixed-semantics target, whereas common-seed adaptive replay does not recover that target. Controlled factorial experiments further re- veal a behavior-by-ranking interaction: controversy-oriented ranking increases participation substantially more under conflict-responsive behavior while redirect- ing activity toward shallower reply structures, a regime we call Shallow Swarm. Mechanism-deletion and execution-trace analyses link the additional volume to conflict-responsive participation and the depth shift to ranking-conditioned expo- sure and reply targeting. The interaction is replicated across independently frozen model-family semantic catalogs and held-out simulator posts, while activation- form and boundary analyses identify settings in which parts of the effect weaken or reverse. Held-out platform evidence supports the conflict-to-volume pathway, whereas evidence for the shallower-depth component remains more limited. Over- all, Rehydration provides a reproducible way to isolate fixed-semantics structural effects from policy-induced semantic adaptation and to test mechanistic claims in generative-agent social simulations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.