Exploring Role-Playing Mechanisms in Large Language Models
Abstract
While large language models exhibit strong role-playing capabilities, the internal mechanisms that support role-conditioned generation remain insufficiently understood. To address this gap, we introduce Contrastive Decomposed Forward Pass for Complete Sequences (CoDePass), which detects role-related computational components by comparing their contributions across matched prompts and complete responses. Our experiments reveal that role playing is organized through a sparse subset of attention heads: recurrent heads retrieve profile and question information, track identity cues, and maintain positional structure, while multi-layer perceptrons (MLPs) progressively transform these signals from explicit role instructions into identity-conditioned response representations. Further behavioral analyses show that while context retrieval, response timing, and baseline semantic construction rely on universally reusable mechanisms across models and characters, identity selection and persona rendering remain highly specific to the assigned role. Finally, targeted fine-tuning of as little as 2% of the model parameters improves role-playing quality over capacity-matched random and full-parameter tuning on held-out characters while preserving general capabilities. These results reveal a division of labor: sparse attention heads control which profile and identity information reaches the response pathway, while MLPs already provide the organized transformation that converts this information into role-related responses. Our code is anonymously available at https://anonymous.4open.science/r/CoDePass-0E45.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.