When Policy Representations Fail: Diagnosis and Repair through Behavioral Evidence
Abstract
Compact policy representations adapt to downstream tasks by optimizing low-dimensional coordinates rather than the full policy parameter vector, reducing adaptation overhead. However, these representations are typically kept fixed after training. When adaptation to a new task fails, expanding the representation is a natural response. We show that such failure is only a symptom, not evidence of insufficient capacity: similar performance degradation can result from limited target-relevant information, inadequate coordinate search, poor organization of behavior within the existing representation, or genuinely missing adaptation directions. We prove that final performance alone is insufficient to uniquely identify the source of failure or determine how the representation should be modified. Because these latent factors are not directly observable, we introduce a behavioral witness: an independently verified successful target adaptation that provides both behavioral and directional evidence. Building on this reference, we develop a witness-guided hierarchical diagnosis and repair framework. Under fixed budgets, the framework tests pre-specified interventions from weaker to stronger and selects the weakest verified repair that satisfies both target-performance and behavior-preservation requirements, expanding the representation only when fixed-capacity alternatives remain insufficient. We evaluate the framework in continuous control and vision-language-action policies. The results show that behavioral evidence supports different levels of sufficient repair, turning default expansion into an evidence-guided, on-demand process that avoids unnecessary representation growth. At the repair stage, the resulting approach also substantially reduces trainable-parameter and peak-memory overhead relative to full fine-tuning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.