SteerSim: De-identifying Developer Behavior in Agent Traces via User Simulation
Abstract
Developer–agent interaction logs are valuable training data, but releasing them poses privacy risks beyond personally identifiable information. We observe that how long developers let an agent run, how closely they scrutinize proposed actions, and how they phrase corrections can form a distinctive behavioral fingerprint that conventional anonymization leaves intact. We introduce SteerSim, an agentic user simulator that replaces the developer in a live coding session to enable privacy-preserving trace synthesis. Operating within the coding agent's environment, SteerSim inspects the repository and takes actions to steer the agent toward task completion. It is conditioned on a persona extracted from real developer sessions, which can be transformed to obscure the original developer's behavioral patterns. We further propose SteerRL, a training algorithm that combines supervised fine-tuning with agentic reinforcement learning using rubric-based rewards. We evaluate both on real sessions from SWE-Chat along three dimensions: human-likeness, re-identification risk, and downstream utility. With Opus 5 as its backbone, SteerSim produces interactions that are nearly indistinguishable from human interactions, as measured by the Turing score. Transforming personas reduces re-identification risk to near chance while preserving downstream utility, achieving a favorable privacy–utility tradeoff. Finally, applying SteerRL to Qwen3.5-4B yields a simulator with performance comparable to Opus 5, offering a practical and low-cost alternative.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.