Beyond Similarity: A PID-Inspired Diagnosis of Skill Composition
Abstract
Modern large language models increasingly rely on retrieving and injecting reusable skills from external libraries. The prevailing paradigm often assumes a positive correlation between skill quantity and performance gain, directly concatenating the top-ranked similar skills into the prompt context. However, our study reveals significant limitations in this approach: aggregating highly related skills frequently leads to severe diminishing returns and information redundancy, or even induces negative interference that degrades model performance. The primary contribution of this paper is the introduction of a novel analytical framework based on Partial Information Decomposition (PID) to systematically investigate the discrepancy between “retrieval similarity” and “actual compositional utility.” By decomposing the model's log-likelihood into shared evidence, skill-specific evidence, and non-additive interactions, this framework enables precise quantitative evaluation of the complementary or interfering effects among multiple skills. Our extensive evaluations across seven math, question-answering, and code generation benchmarks reveal that top-ranked skill combinations typically suffer from information redundancy, yielding a marginal gain that is frequently negative relative to the best single skill. Furthermore, as an extension of this analytical framework, we conduct a preliminary exploration of using PID as a guiding strategy for skill selection; empirical results indicate that filtering skills via PID can mitigate inter-skill interference to some extent, offering a promising pathway for optimizing multi-skill composition in the future.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.