DHKR: Data-Demanded Capacity Curricula for Hierarchical Representation Routing in Frozen Language Models
Abstract
While adapting language models often relies on rigid capacities and static routing structures, we introduce Dynamic Hierarchical Knowledge Routing (DHKR), a general framework for representation adaptation on frozen backbones. DHKR treats structural intervention availability as a control variable distinct from input-conditioned allocation and parameter optimization. Using a Knowledge Steering Layer (KSL) as an external portfolio of nonlinear representation interventions, DHKR uses a rule-based controller to retain or release paths within a configured intervention portfolio. Counterfactual analyses show that replacing learned weights with uniform or averaged variants increases validation loss, while restricted availability yields lower validation loss than full-from-start baselines under a matched regularized protocol. These findings show that allocation, structural availability, and Controller execution produce distinct outcomes in the tested setting. Separate trajectories verify that training continues after portfolio release.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.