acceptodds
Under review as a conference paper at ICLR 2027

Can Circuit Analysis Predict Capability Transfer Across Tasks and Modalities?

Abstract

Circuit analysis is mainly used post-hoc to explain model behavior, but could it also be used to guide training decisions? We show that a model’s internal computation, measured before fine-tuning, can predict how training on one task will affect another. Across four vision-language models from two families and 18 tasks with textual and visual variants, we establish three findings. First, task circuits present in the base model largely persist after fine-tuning, meaning that circuits discovered before training remain representative afterward. Second, relationships between these base-model circuits (measured by cross-task circuit accuracy) can predict capability transfer: our circuit-based measure correlates with cross-task fine-tuning gains up to r = 0.77, and is competitive with an established black-box transferability estimation method while capturing non-trivial transfer across task types. Third, we characterize when these predictions succeed and fail, and find they are most reliable when task circuits are well localized, but break down when similar-looking circuits conceal conflicting computations. Together, these findings show how circuits measured once can inform capability transfer, turning circuit analysis from an explanatory tool into a predictive one.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.