Behavior Amplification: Discovering and Steering Model-Specific Reasoning Strengths in LLMs
Abstract
Recent work has advanced the ability of Large Language Models (LLMs) to reason about complex tasks. However, the majority of methodologies treat reasoning as a monolithic capability, whereas state-of-the-art models exhibit a high degree of diversity in their underlying reasoning behaviors. In this paper, we study how model-specific behavioral patterns are associated with success or failure in question answering. We propose a data-driven framework for discovering model-specific Behavioral Strength Sets: reasoning behaviors that are strongly correlated with answer correctness. These correlated behaviors are identified and selected via contrastive analysis and significance testing. Subsequently, we introduce Inherent Strength Amplification (ISA), a steering framework that leverages these behaviors to construct targeted interventions, including prompt injection, behavioral reflection, supervised fine-tuning (SFT), and reinforcement learning (RL). Across three model families and three question answering benchmarks, ISA consistently improves answer correctness over unsteered baselines, with gains of up to 41.0 EM and 29.2 F1 points. Our findings indicate that the effectiveness of each intervention depends upon the behaviors being amplified: behaviors that can be directly expressed as inference-time instructions are often improved through prompting or reflection, whereas behaviors requiring consistent multi-step execution derive greater benefit from parameter-level training. These outcomes demonstrate that effective reasoning improvement requires the alignment of interventions with a model’s specific behavioral profile, rather than the application of a universal optimization strategy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.